Ask the vendor how they match your product to a competitor's listing, how often that match is wrong, and how you find out when it is. Coverage numbers, refresh rates, dashboards and repricing rules are all downstream of matching.
A demo that shows twenty perfectly matched products proves nothing about your catalogue. Insist on a pilot against your own SKUs, and check the matches by hand yourself.
Why matching is the question, not scraping
Fetching a page is a solved problem you can buy from several suppliers. Deciding that your article 88213 is the same thing as their listing, at the same pack size, the same vintage, the same voltage, the 25 kg bag rather than the 5 kg bag, is not solved and never fully will be.
A wrong match does not look wrong on a dashboard. It looks like a competitor who is 22% cheaper. Somebody then cuts a price to chase a product that is not the same product, and the margin leaves quietly.
This is why so many pricing mistakes end up being prices set too low. Bad matches almost always pull you downwards: a smaller pack, an older model, a grey import or a bare unit without accessories always looks cheaper than the thing you actually sell.
Questions about matching
- How do you match: identifier, text similarity, a model, human review, or a combination? Ask for the order of precedence.
- What is your precision on my catalogue, measured on a random sample that I choose rather than one you choose?
- Who reviews low-confidence matches, and is that review inside the quoted price or billed separately?
- Do I see a confidence score per match, and can I filter and sort on it?
- Can I reject a match, and does the rejection survive the next crawl and the next relisting?
- Do you treat variants and pack sizes as distinct products, or fold them into one parent?
- What happens to a match when the competitor changes the URL, or delists and relists the item?
Questions about freshness and coverage
Coverage gets quoted three different ways and only one of them is useful. Ask for the share of your SKUs that had at least one live competitor price in the last 24 hours. Everything else is a bigger number that means less.
- What share of my SKUs have a live competitor price today, not ever?
- Which of my named competitors are blocked to you right now, and what is the plan for them?
- How often is each competitor refreshed, and is that number a target or a measurement?
- Do I get a timestamp and a clickable source URL on every price?
- When a fetch fails, do you show the last known price or nothing? Ask them to use the word "stale" out loud.
- How do you detect a parser that has silently started reading the wrong element on the page?
The awkward parts of a price
These decide whether the number is comparable at all, and they are where cheap tools quietly fall apart:
- VAT: gross or net, and whose rate on a cross-border listing.
- Delivery: included, excluded, or free above a threshold. A 4.90 euro delivery charge on a 19 euro product is about 26% of the item price, since 4.90 divided by 19 is 0.258.
- Deposits and levies: Pfand, WEEE fees, tyre and battery charges.
- Multi-packs and unit price. A six-pack at 54 euro is 9 euro a bottle, and those two figures must never share a column.
- Stock: an out-of-stock listing is not a competing price and should not trigger a rule.
- Marketplaces: whose price is being captured, the buy box winner's, the cheapest seller's, or the marketplace's own?
- Promotions: is the strike-through price stored separately from the effective price?
How to run a pilot that proves something
Two weeks, your catalogue, your competitors, and one afternoon of your own attention.
- Pick 100 SKUs at random from live stock, not your best sellers and not your cleanest data.
- Give the vendor the export you would give in production, without hand-cleaning it first.
- After the first crawl, open the matches for 50 of those SKUs and click every source URL yourself.
- Count the wrong ones. Precision is correct divided by checked.
- Reject the bad matches, then repeat the count in week two to see whether the rejections stuck.
Do the arithmetic out loud. If 6 of 50 checked matches are wrong, precision is 44 divided by 50, which is 88%. On a catalogue of 10,000 SKUs with 4 competitors each, that is 40,000 pairs, and 12% of them wrong is 4,800 bad comparisons. If your rules act automatically whenever a competitor looks cheaper, several thousand automated price cuts are now standing on bad data.
88% precision is not a scandal. It is normal for a first pass. What matters is whether the vendor knows the number, tells you the number, and has a review path for the remaining 12%.
Questions about leaving
- Can I export every match, with source URLs and confidence scores, as CSV and through an API?
- Do I own the match decisions and corrections my team made, or do they stay with you?
- What is the notice period, and does the contract auto-renew?
- How does the price move when my SKU count or competitor count grows, and is there a cap?
- Is API access included, and what are the rate limits?
- How long is my data retained after termination, and can I have it deleted on request?
What a good answer sounds like
Vendors who do this properly answer in numbers and hedges. They name the competitors they currently cannot fetch. They quote precision with a sample size attached. They show you a match with a confidence of 0.61 and explain why a person looked at it before it reached your dashboard.
Vendors who do not, answer with adjectives, promise complete coverage, and demonstrate on their own catalogue rather than yours.
PriceRoom is built on the assumption that the match is the product and the crawl is plumbing, which is also why we would rather you check 50 matches by hand during a pilot than watch a dashboard move. A tracker that tells you where you are already the cheapest is easy to build and worth very little. The value is in the places where you were below the market and nobody noticed.