It's always worth showing properly designed and conducted trials--though it's bad for pageviews. Nonetheless, SNAP authors show the way to answering questions in biomedicine
1/ A real-world "OR ≠ RR" lesson from the SNAP trial (NEJM 2026): cefazolin vs antistaphylococcal penicillin for MSSA bacteremia. Small slip, big teaching point. 🧵
2/ 90-day mortality: 15.0% (cefazolin) vs 17.0% (penicillin). Adjusted OR 0.81.
Tempting reading: "a 19% reduction in death" (= 1 − 0.81).
Catch: that's a reduction in odds, not in deaths.
3/ The actual relative risk = 15.0 / 17.0 = 0.886 → relative reduction ≈ 11%.
Absolute reduction ≈ 1.9 points.
When the outcome is common (here 15–17%, well above 10%), the OR overstates the effect vs the RR.
4/ Bigger nuance: mortality superiority isn't even established.
95% credible interval 0.59–1.12 → crosses 1. Probability of superiority 89.8%.
Cefazolin is non-inferior on mortality; "19% fewer deaths" assumes both a magnitude and a certainty the data don't support.
5/ Rule of thumb: OR ≈ RR only for rare events (<10%). Above that, ALWAYS convert OR to RR and to an absolute difference before you communicate.
SNAP's robust signal was safety (less AKI, fewer discontinuations), not mortality
What makes SNAP worth holding up is not just the result, it is that the design did the thinking before the data arrived. The adaptive Bayesian structure with a pre-specified stopping rule is the antidote to the two failure modes that wreck most trials, the underpowered study that limps to a null and the overpowered one that chases a fragile p-value long past clinical meaning.
Stopping near a fifth of the planned enrollment because the question was actually answered is the discipline most trials lack. From the data side I would add that this kind of design also resists the comparator and endpoint games, because the primary endpoint and the decision thresholds are locked in advance.
Do you think the main barrier to more SNAP-style trials is statistical literacy among investigators, or the incentive structure that still rewards the large fixed-design trial?
I understand the importance of trials, but not being a medically informed reader, why would the profession not conclude after the trial to use both drugs at the same time?
My ignorance: how is 15% v 17% a 19-20% reduction in mortality? (sorry don't remember which number was quoted). Eyeballing it, it looks more like an 11% reduction in mortality (2% of 17%). Clearly my stats knowledge is woefully deficient! I'll own my stupidity, if someone can tell me how the figure is meant to be calculated.
You're not being ignorant — your instinct is exactly right, and your ~11% is the correct number.
The 19–20% came from the odds ratio (0.81): 1 − 0.81 = 19%. But that's a reduction in odds, not in deaths.
What you eyeballed is the relative risk: 15/17 = 0.88 → ~12% relative reduction (≈2 absolute points). That's the clinically honest figure.
Why the gap? OR ≈ RR only when the outcome is rare (<10%). At 15–17%, the OR drifts further from 1 and overstates the effect. Whenever you see an OR for a common outcome, convert it: RR = OR / [(1−p₀) + p₀·OR]. Here: 0.81 / [0.83 + 0.17·0.81] ≈ 0.84.
Trusting your "2% of 17%" gut was the smart move. 👏
Thanks for the highlighting John! I keep reminding our SNAP team that this is why we do this work - to improve outcomes for patients. Being praised for doing good EBM is an absolute bonus!!
Outstanding study and outstanding commentary. I only take small issue with
Here are the numbers for the protocol-adherent: OR 0.88, 95% credible interval 0.61 to 1.26.
So, technically, the upper-bound of the 95% confidence interval of 1.26 is greater than the
margin of 1.20. And the posterior probability of noninferiority in this population drops to 95.4%,
well below the 99% stopping threshold that was used for the primary analysis.
Don't look at the credible interval (which I hope is really a highest posterior density uncertainty interval); look only at the posterior probability of non-inferiority. Then your concern should not be with evidence for non-inferiority among adherers but rather with the use of an arbitrary evidence threshold such as 0.99. P(NI) = 0.954 in a subset analysis is quite impressive.
Minor quibble: Probabilities are between 0 and 1 so all of the probabilities you are referencing should actually be labeled as 100x probabilities. Best to stick with actual probabilities which are also less likely to be misunderstood.
Agree 100%! This is the kind of study we should be doing, not just in ID, but all of medicine. Note that the USA did not participate -- this is because of the very high cost of running trials here compared with other countries, an unfortunate reality.
OR ≠ RR when the outcome is common — a 30-second lesson.
"Risk" = events / everyone. "Odds" = events / non-events.
When events are frequent (>10%), the OR drifts further from 1 than the RR → it overstates the effect.
SNAP (NEJM 2026): mortality 15% vs 17%, OR 0.81. "19% fewer deaths" is the odds reduction. True RR ≈ 0.89 (~11%).
Always convert OR → RR + absolute difference before you communicate
1/ A real-world "OR ≠ RR" lesson from the SNAP trial (NEJM 2026): cefazolin vs antistaphylococcal penicillin for MSSA bacteremia. Small slip, big teaching point. 🧵
2/ 90-day mortality: 15.0% (cefazolin) vs 17.0% (penicillin). Adjusted OR 0.81.
Tempting reading: "a 19% reduction in death" (= 1 − 0.81).
Catch: that's a reduction in odds, not in deaths.
3/ The actual relative risk = 15.0 / 17.0 = 0.886 → relative reduction ≈ 11%.
Absolute reduction ≈ 1.9 points.
When the outcome is common (here 15–17%, well above 10%), the OR overstates the effect vs the RR.
4/ Bigger nuance: mortality superiority isn't even established.
95% credible interval 0.59–1.12 → crosses 1. Probability of superiority 89.8%.
Cefazolin is non-inferior on mortality; "19% fewer deaths" assumes both a magnitude and a certainty the data don't support.
5/ Rule of thumb: OR ≈ RR only for rare events (<10%). Above that, ALWAYS convert OR to RR and to an absolute difference before you communicate.
SNAP's robust signal was safety (less AKI, fewer discontinuations), not mortality
What makes SNAP worth holding up is not just the result, it is that the design did the thinking before the data arrived. The adaptive Bayesian structure with a pre-specified stopping rule is the antidote to the two failure modes that wreck most trials, the underpowered study that limps to a null and the overpowered one that chases a fragile p-value long past clinical meaning.
Stopping near a fifth of the planned enrollment because the question was actually answered is the discipline most trials lack. From the data side I would add that this kind of design also resists the comparator and endpoint games, because the primary endpoint and the decision thresholds are locked in advance.
Do you think the main barrier to more SNAP-style trials is statistical literacy among investigators, or the incentive structure that still rewards the large fixed-design trial?
I understand the importance of trials, but not being a medically informed reader, why would the profession not conclude after the trial to use both drugs at the same time?
My ignorance: how is 15% v 17% a 19-20% reduction in mortality? (sorry don't remember which number was quoted). Eyeballing it, it looks more like an 11% reduction in mortality (2% of 17%). Clearly my stats knowledge is woefully deficient! I'll own my stupidity, if someone can tell me how the figure is meant to be calculated.
You're not being ignorant — your instinct is exactly right, and your ~11% is the correct number.
The 19–20% came from the odds ratio (0.81): 1 − 0.81 = 19%. But that's a reduction in odds, not in deaths.
What you eyeballed is the relative risk: 15/17 = 0.88 → ~12% relative reduction (≈2 absolute points). That's the clinically honest figure.
Why the gap? OR ≈ RR only when the outcome is rare (<10%). At 15–17%, the OR drifts further from 1 and overstates the effect. Whenever you see an OR for a common outcome, convert it: RR = OR / [(1−p₀) + p₀·OR]. Here: 0.81 / [0.83 + 0.17·0.81] ≈ 0.84.
Trusting your "2% of 17%" gut was the smart move. 👏
A good study like this one always shines a light and a bad study always wallows in discontent.
Thanks for the highlighting John! I keep reminding our SNAP team that this is why we do this work - to improve outcomes for patients. Being praised for doing good EBM is an absolute bonus!!
I'm very proud to have contributed to this trial and learnt a great deal from it's leaders.
Very interesting. Thank you for the summary!
Great news, but what would be used in PCN Allergic with first generation cephalosporin allergies?
Outstanding study and outstanding commentary. I only take small issue with
Here are the numbers for the protocol-adherent: OR 0.88, 95% credible interval 0.61 to 1.26.
So, technically, the upper-bound of the 95% confidence interval of 1.26 is greater than the
margin of 1.20. And the posterior probability of noninferiority in this population drops to 95.4%,
well below the 99% stopping threshold that was used for the primary analysis.
Don't look at the credible interval (which I hope is really a highest posterior density uncertainty interval); look only at the posterior probability of non-inferiority. Then your concern should not be with evidence for non-inferiority among adherers but rather with the use of an arbitrary evidence threshold such as 0.99. P(NI) = 0.954 in a subset analysis is quite impressive.
Minor quibble: Probabilities are between 0 and 1 so all of the probabilities you are referencing should actually be labeled as 100x probabilities. Best to stick with actual probabilities which are also less likely to be misunderstood.
Agree 100%! This is the kind of study we should be doing, not just in ID, but all of medicine. Note that the USA did not participate -- this is because of the very high cost of running trials here compared with other countries, an unfortunate reality.