How We Measure: Protocols, Sample Sizes and Dates

Use AI to summarise this article

Every number on this site is supposed to be repeatable by you, which means the method has to be published before the result is. This page is that method: the test dataset and how to rebuild it exactly, the rules the timers run under, how a price gets read and dated, what we count as a measurement and what we count only as research, and the honest list of the things we have not run yet. If a page here ever states a figure whose protocol is not described below, that is a defect on our side and it should be reported.

Last updated 27 August 2026. This page is revised whenever a protocol changes or a run completes.

What is this page for?

It is the reference every measured figure on this site points back to. Rather than repeat the sample size, the date, the timer rules and the caveats on every article, they are stated once here in full. A reader who wants to check our arithmetic, repeat a run, or decide how much weight a number deserves starts on this page.

What counts as a measurement, and what counts as research?

We use two labels and we never blur them.

Measurement. We did the thing, we recorded it while doing it, and the record exists. A migration we ran. A setup we timed. An invoice we paid. A cap we hit. A cancellation we requested and got an answer to. A measured page states the sample size, the date and the procedure above the conclusion, and it points back here.

Research. We read the documentation, the terms or the pricing page and reported what it said. That is useful, and it is often the only checkable thing published on a subject, but it is not a measurement. A research page says so in the same size type as its conclusion, and it never implies a test happened.

The distinction is not decoration. Almost every page that ranks for these queries describes an outcome without ever saying whether anyone performed it. Where we have not performed it, this site says so on the page rather than in a disclaimer at the foot.

What is in the test dataset?

A synthetic 500-record fixture, not a live business. It stands in for a small company's CRM in every hands-on run, so that each run starts from identical data and results can be compared across vendors rather than across datasets.

File Rows Object it stands for
companies.csv 100 Company, Organization or Account
contacts.csv 500 Contact, Person or Lead
deals.csv 62 Deal or Opportunity, across two pipelines
activities.csv 228 Call, email, meeting or task
notes.csv 99 Free-text note

It is generated deterministically by a script from the seed 20260819, and its manifest carries a SHA-256 hash for each of the five files. Rebuild it with the same seed and you get byte-identical files, which means any run described on this site can be repeated on your own account against exactly the data we used.

No real person, company, email address, phone number or deal appears in it. Domains are reserved .example addresses.

Diagram of the test fixture: companies.csv 100 rows, contacts.csv 500, deals.csv 62, activities.csv 228 and notes.csv 99, followed by the eleven deliberate data traps built into it, from six phone formats to notes over 500 characters.
Rebuild it with seed 20260819 and you get byte-identical files, which is what makes any run described here repeatable on your own account. Built, not yet run.

The traps in it are deliberate

A clean dataset proves nothing, because nobody's CRM is clean. The fixture is built with the failures a real import has to survive already planted in it:

  • Six phone formats, including parenthesized US, dashed US, international with a country code, bare digits, a UK mobile, and blank.
  • Duplicate emails: roughly 3 percent exact duplicates and roughly 2 percent case-variant duplicates, where Name@ and name@ are one address to a human and often two records to an importer. Roughly 5 percent of contacts carry no email at all.
  • Blank fields that look required: job titles, employee counts, revenue, lifecycle stages and deal amounts are empty on some rows.
  • Custom fields on every object, including a date field, a numeric score, a boolean-ish text field and picklists, because custom fields are where most migrations quietly lose data.
  • Four currencies on the deals: USD, EUR, GBP and CAD.
  • Two pipelines with different stage sets, which is the single most common reason a deal import lands everything in stage one.
  • Multi-value tags inside one cell, separated by semicolons.
  • Notes over 500 characters, and datetimes carrying a time component rather than a bare date.
  • Non-US addresses and blank state fields.

Every unmapped field, dropped association, truncated note and rejected row in a run is counted against these files, so the count means the same thing every time.

How is a price read, and what does "checked on" mean?

A price on this site is a record of what one vendor published on one page on one named day. It is not a forecast, and it is not a promise about the day you read it.

The rules the price reads run under:

  1. Read from a US-located browser, logged out, in a private window, with the currency selector set to USD where the page offers one. A price a vendor serves in another currency is recorded in that currency and never converted into a dollar figure the vendor does not charge. Where a conversion is genuinely useful, the rate and its date sit next to it and the original figure stays visible.
  2. Toggle both ways. Nearly every pricing page defaults to the position showing the lower number. We record the monthly figure and the annual figure, and the two-year figure where one exists, and we name the period on the same line as the number.
  3. Record the page URL and the moment of reading, and put the date on the page you are reading.
  4. Record seat minimums and seat buckets from the pricing page or from the vendor's support documentation, and say which of the two it came from.
  5. Never edit a past reading. A correction is a new dated record that says which reading it corrects. That is what makes a price history a history rather than a rolling opinion.
  6. Repeat about every 30 days. A missed month is recorded as missed rather than backfilled, because a backfilled snapshot is a guess wearing a date.

Some vendors cannot be read this way, and we say which. Automated requests to some pricing pages are refused outright. Some pages build their price table with JavaScript in a way that returns empty to an automated read. Some serve a local currency based on where the reader is sitting. Those vendors are read by hand, or they are left out of the figure and named as an exclusion. A missing vendor is a finding. A vendor filled in from a comparison article is a fabrication.

What is never a source

Numbers seen in a search result snippet, a ranking article, a directory listing or a generated summary are not sources here, ever. Only the vendor's own page, the vendor's own documentation, or our own record.

That is not pedantry. When we read the live results for these queries on 19 August 2026, two dated 2026 articles about the same product disagreed with each other about the price of the same plan, and neither showed a dated screenshot. Whichever of them was right, copying either one would have been guessing.

How are the timed runs timed?

Every timed task uses the same task list across every product, so the comparison is between the products rather than between our moods on two different days.

  • The clock starts at the first click after account creation and stops when the stated task is complete and visible in the interface, not when it was submitted.
  • Each step is timed separately, so a slow import cannot hide a fast setup.
  • Stalls are recorded, not smoothed. Where a step could not be completed at all on the tier being tested, it is logged as blocked with the reason rather than dropped out of an average.
  • The sample size is stated as the number of systems and the number of runs, never as "typically" or "usually".

What have we not tested?

This is the part most pages leave out, so it goes here rather than in a footnote.

As of 27 August 2026, no hands-on run on this site has been completed. Specifically:

  • No migration has been run. The method is written and the fixture is built, but no data has been moved between two CRMs by us. Pages in the migration cluster are therefore research versions, and each one says so on the page itself.
  • No setup stopwatch has been run. No product has been timed through the setup task list.
  • No paid seat has been bought by us at any vendor. No invoice exists.
  • No calls have been made through any CRM's dialer, so no cost per connected call has been computed from a bill.
  • No free-plan cap has been hit. Where we quote a free-tier limit, it is the vendor's stated cap, not an observed ceiling. Those are different numbers often enough that we treat the gap between them as a subject rather than an assumption.
  • No account has been canceled, so nothing here describes a cancellation, a downgrade or a refund from experience.

What we do have today is the fixture, the protocols on this page, and dated reads of vendors' published pricing pages. That supports the cost pages, and it does not support anything else, so the cost pages are what exists.

What this page does not do

It does not tell you which CRM to buy, and it does not rank them. It does not assess whether a product is good, only what it published and, once the runs exist, what it did when we used it. It cannot tell you what you will pay, because your seat count, your record count, your contract and any negotiated discount all sit outside a published price. And the protocols above bind us, not the vendors: a vendor is free to change a price, a cap or a tier the day after we read it, which is precisely why every figure carries a date.

FAQ

Can I repeat your runs myself?
Yes, and that is the point of publishing the seed and the hashes. Rebuild the fixture with seed 20260819, check the five SHA-256 hashes against the manifest, and you are starting from data identical to ours. The task lists and the timer rules are described above.

Why use synthetic data instead of a real company's CRM?
Because a real export cannot be published, and a run nobody can repeat is an anecdote. The fixture is declared as synthetic on every page that uses it. Its edge cases are drawn from the failures real imports hit, and they are listed above so you can judge for yourself whether they match your own data.

Why do some vendors have no price on your pages?
Because their pricing page could not be read under the rules above on the day of the read, or because the tier is quote-only. We name the vendor and the reason rather than filling the gap from somewhere else. A tier with no published price is itself a finding worth stating.

How often are prices rechecked?
About every 30 days, with each new reading added as a new dated record rather than overwriting the old one. Any page carrying a price shows the date it was read. If that date is more than a month old, treat the figure as historical.

What happens to a research page once you have actually run the thing?
It is replaced by the measured version, and the change is dated. The research version's conclusion is not quietly edited to match the result, because the gap between what the documentation implied and what actually happened is usually the most useful thing on the page.

Get help