Cold email benchmarks: open and reply rates
Every cold email benchmark you have read was published by a company that sells cold email software. That does not make the numbers fake, but it does make them useless as a target. Vendors report on their own customer base, they count opens and replies differently, they quietly exclude the accounts that churned, and they pick the year that flatters the chart. Two "industry averages" can differ by a factor of five and both be technically true.
So this article does something less satisfying and more useful. It gives you order-of-magnitude expectations instead of decimal points, explains why one of the two headline metrics is now broken beyond repair, and shows you how to build your own baseline over a few batches — because your baseline is the only number that can tell you whether today's campaign is working.
Why published benchmarks fall apart under inspection
Three problems make cross-vendor benchmarks incomparable:
- Selection bias. The dataset is whoever pays that vendor. If a tool is popular with recruiters, its "average" reflects recruiting outreach. If it is popular with agencies blasting scraped lists, the average collapses. Neither describes you.
- Incompatible definitions. Is a reply an out-of-office? A "remove me"? A bounce that lands in the inbox? Some tools count any inbound message, some count only human replies, some count only positive ones. The same campaign can honestly be reported at 2% or 6%.
- Denominator games. Reply rate per contact and reply rate per email sent differ by the number of follow-ups. A five-step sequence sending to 200 people has sent 1,000 emails. Divide by the wrong number and you double or halve the result.
When someone quotes you a precise figure — "the average B2B reply rate is 8.4%" — the correct response is to ask what counted as a reply and who was in the sample. You will rarely get an answer.
Open rate is now a broken metric
Open tracking works by embedding an invisible image; when the recipient's client loads it, the sender records an open. Apple's Mail Privacy Protection changed that. Apple Mail now prefetches images through a proxy for many users, whether or not a human ever looked at the message. Corporate security gateways and link scanners do the same thing, and some of them also click every link to check it for malware.
The consequences:
- Your open rate is inflated by an unknown, campaign-specific amount. It is not a constant you can subtract.
- A reported jump from 35% to 50% may reflect a change in the mail clients on your list, not your subject line.
- Click-through can be inflated by security scanners too, which is why a "click" from a corporate domain seconds after delivery is usually a robot.
Practical rule: reported open rates for cold B2B campaigns commonly land somewhere between 20% and 60%, and anywhere in that band tells you almost nothing about human interest. Treat open rate as a coarse deliverability smoke alarm, not a performance metric. If it drops from 40% to 5% overnight, you have a delivery problem. If it moves from 38% to 44%, you have noise. Never A/B test subject lines on opens alone; test them on replies, which requires more volume and more patience.
Realistic ranges for cold outbound
Everything below is an order-of-magnitude expectation for genuinely cold B2B outreach — people who have never heard of you — measured per contact, across a full sequence including follow-ups. These are not measured facts about your market; they are the bands within which most honest campaigns sit.
Reply rate
- 1–5% total replies is the normal range for a decent list and a specific message. Most campaigns live near the bottom of it.
- 5–10% usually indicates a narrow, well-researched list, a strong relevance signal in the first line, or an offer that solves an urgent and obvious problem.
- Above 10% almost always means the list was not really cold — existing contacts, event attendees, inbound signals, referrals — or that the sample is tiny. It is very rarely a template effect. Anyone selling you a template that "gets 30% replies" is describing a warm list.
- Positive replies — people who want to talk, not those declining — typically run at roughly a third to a half of total replies. Plan on 0.5–2% positive replies per contact.
- Meetings booked land in the 0.2–1% range per contact for most B2B offers. At the top of that band, 1,000 well-chosen contacts produce something like ten conversations.
Do the arithmetic before the campaign, not after. If your product needs 20 demos a month and your realistic booking rate is 0.5%, you need roughly 4,000 contacts a month — which for a two-person team is a different business than the one you thought you were running. That calculation kills more bad outbound plans than any benchmark article ever will.
Where industry and region actually shift the numbers
Vertical-by-vertical benchmark tables imply a precision nobody has. What is defensible is the direction of the effect:
- Saturation beats industry. Roles that receive dozens of pitches a day — VP Sales, marketing leadership, IT decision-makers at mid-size and larger firms — reply less, often at the bottom of the range or below. Owner-operators of local businesses, trades, clinics and small manufacturers reply far more, because almost nobody writes to them thoughtfully.
- Company size moves inversely to reply rate. Writing to a 6-person studio in Toronto usually beats writing to a 6,000-person enterprise, where your message must survive gatekeepers, procurement and a spam filter tuned by a security team.
- Regional channel preference matters more than regional "email culture". In much of North America and Northern Europe, email is still the default business channel. In parts of Southern Europe, Latin America, the Middle East and South and Southeast Asia, a business phone number is often a messaging account, and a short, polite message there outperforms an email that will never be opened. The same list can produce a 1% email reply rate and a materially higher messaging reply rate.
- Language. Writing in the recipient's language reliably outperforms English-by-default outside English-speaking markets. This is one of the few "tricks" with a real, repeatable effect.
Deliverability numbers that are not opinions
Unlike reply rates, these have hard thresholds because mailbox providers enforce them:
- Bounce rate under 2% is healthy. Above 5% means your list is stale or unverified and you are damaging your sending domain. Above 10%, stop the campaign today and re-verify.
- Spam complaints above 0.1% — one in a thousand — put you in trouble with major providers. There is no safe way to run consistently above it.
- Volume per mailbox: 20–50 cold sends per day per address is the range most senders can sustain. Higher volumes need more mailboxes and warmed domains, not a bigger daily number on one account.
Build your own baseline in three or four batches
Your market, your offer and your list quality dominate any published average. So stop comparing and start measuring. The procedure is boring and it works.
- Fix the variables. One segment, one offer, one sequence structure. Do not change the ICP and the copy at the same time — you will learn nothing.
- Send in batches of 150–250 contacts. Small enough to stop cheaply, large enough that the result is not pure luck.
- Run three or four batches before drawing any conclusion. Somewhere between 600 and 1,000 contacts, your reply rate stops jumping around and settles into a range. That range is your baseline.
- Record four numbers per batch: delivered, total replies, positive replies, meetings. Opens go in a separate column labelled "unreliable", if you record them at all.
- Only then optimise. Change one thing per batch and compare against the baseline, not against a blog post.
Be honest about statistics. At a true reply rate of 1%, sending 200 emails produces zero replies about 13% of the time by chance alone. At 500 sends the odds of a genuine 1% campaign showing zero replies fall below 1%. This is why a single 200-contact batch cannot tell you whether your message works, and why "we tested it, it didn't work" after one small batch is usually not a test at all. Detecting the difference between a 1.5% and a 3% campaign reliably takes thousands of sends, which is why most small teams should chase list quality — a change that moves the number severalfold — rather than copy tweaks that move it fractionally.
Diagnostics: what your numbers are telling you
When results disappoint, the fault is in one of three places, and the numbers separate them cleanly.
Zero replies after 200 sends
Not yet alarming on its own — see the arithmetic above. Check delivery first: if your reported open rate is also near zero and no bounces came back, your mail is probably landing in spam and being silently discarded. Fix authentication and warm-up before writing another word of copy. If opens look normal at 200 sends with zero replies, continue to 400–500 and then judge.
Zero replies after 500–600 sends
Now it is a real signal. In descending order of likelihood: the list is wrong (people who cannot buy or do not have the problem), the message is about you rather than them, or the offer asks for too much too early. Fix the list first. A mediocre message to the right 200 people beats a brilliant message to the wrong 2,000.
Replies exist but are all negative or "wrong person"
This is a targeting failure, and it is good news, because it is cheap to fix. Repeated "not my department" replies mean your role filter is wrong. Repeated "we already use X" means you are targeting a mature segment and need a displacement angle or a different segment. Repeated "how did you get my details" means your opening line failed to establish a reason for contact.
Bounce rate above a few percent
A list-hygiene problem, not a copy problem, and it is urgent because it damages your domain. Verify addresses before sending, drop catch-all domains you cannot confirm, and never import a list you did not build or inspect. Bounces above 5% with a fresh list usually mean the source is scraped, old, or invented by pattern-guessing.
Opens fine, clicks fine, replies zero
Classic ask-mismatch. Cold recipients rarely book a 30-minute call with a stranger. Ask a question that costs ten seconds to answer instead of requesting time. Also check whether your "clicks" are security scanners — a click within seconds of delivery from a corporate domain, with no dwell, is a machine.
Everything was fine, then dropped by half
Suspect delivery before creativity: a sudden volume increase, a new sending domain without warm-up, a link to a low-reputation domain, or a segment full of invalid addresses. Reply rates decay gradually as you exhaust a market; they do not halve overnight because your writing got worse.
The metrics worth watching instead
If you drop open rate to a smoke alarm, what replaces it at the top of the dashboard?
- Positive reply rate per contact — the only rate that correlates with revenue.
- Contacts per meeting — the number that tells you how much list you need to hit a target. Every other metric is downstream of this one.
- Cost per positive reply — data, tools and your hourly rate, divided by conversations. Compare this against paid channels honestly; outbound often wins on cost and loses on scale.
- Bounce and complaint rates — hard limits, checked every batch.
Track follow-ups separately as well. Across most sequences, a meaningful share of replies — often a third to a half — arrives after the first message, which is why sending one email and concluding "cold email doesn't work" is a measurement error rather than a finding. Two or three follow-ups is the useful range; beyond that you mostly generate complaints.
What to do with all of this on Monday
Pick one segment you can describe in a sentence — "independent dental clinics in Toronto with their own website" is a segment; "SMBs in North America" is not. Build a clean list of 200 of them with a verified contact channel each, which is exactly the job JustLeadIt was built for: you describe the niche and the city, and you get contacts with emails, phones and social profiles rather than a spreadsheet you still have to research.
Send batch one. Record delivered, replies, positive replies, meetings. Repeat three more times, changing nothing. In a month you will own a baseline that no benchmark article can give you, and every future decision — new copy, new segment, new channel — gets measured against a number that is actually about your business. That is worth more than knowing what the average is for somebody else.