The Silent Question: How do we quantify benefit?
Not long ago, I sat in a project meeting reviewing the risk analysis for a new medical device.
Every potential harm had been diligently identified and scored for probability and severity in a detailed FMEA.
But when I asked, “So... how are we quantifying the benefit in our benefit-risk assessment?”, the room fell silent.
A colleague finally ventured, “We list the benefits in the IFU, I guess.” In other words, while risks got rigorous tables and ratings, the device’s benefits were handled with a vague one-liner (“improves patient outcomes, trust us”).
Sound familiar? In many development teams, risk analysis is a quantitative exercise, yet benefit analysis remains a qualitative afterthought.
The result is often a pro-forma statement in regulatory documents: “Based on expert judgment, the benefits outweigh the risks.”
Without evidence or numbers, however, that statement can sound like regulatory boilerplate rather than a credible conclusion. How can we make our benefit-risk determinations more concrete, logical, and convincing – perhaps even quantifiable – instead of just repeating regulatory karaoke for the hundredth time?
In this article, we’ll explore what my deep dive into this topic revealed and how I approach benefit-risk assessments for medical devices by combining qualitative insight with quantitative rigor.
We’ll also incorporate useful metrics like NNT/NNH and emerging best practices (including the new French XP S99-223 standard) that can help put real numbers behind the “benefit” side of the equation.
✍️ Note: This article was heavily inspired by the excellent 2022 white paper from RQM+ titled "A Quantitative Approach to Benefit-Risk Determination" by Bethany Chung and Jaishankar Kutty. Their work helped shape my understanding of the benefit-risk framework, and I’ve expanded on it here by integrating methods like NNT/NNH, the XP S99-223 standard, and practical case applications based on my consulting experience.
Why Benefit-Risk Analysis Needs More Than a Statement
Regulators today expect more than a generic assurance that benefits outweigh risks – they expect evidence.
Under the EU Medical Device Regulation (MDR 2017/745), a device cannot be placed on the market unless you demonstrate that its benefits outweigh its risks, ideally in a “clearly quantified and documented” analysis.
In fact, MDR explicitly defines benefit as a positive impact on health expressed in meaningful, measurable, patient-relevant clinical outcomes.
This means manufacturers can’t get away with vague claims like “improves quality of life” without data – you need to specify how it improves health, how much, and how you know. Similarly, ISO 14971 (the global risk management standard for devices) requires that any residual risks be acceptable in light of the clinical benefits, essentially mandating a benefit-risk analysis as part of risk management.
Crucially, neither MDR nor ISO 14971 hands us a formula for determining an “acceptable” benefit-risk trade-off. The MDR doesn’t define a specific ratio or threshold, nor exactly how to justify that benefits trump the remaining risks.
Lacking concrete guidance, many companies have defaulted to purely qualitative arguments in their Clinical Evaluation Reports and risk management files.
All too often, this results in subjective assertions that might not satisfy a skeptical reviewer – after all, two people could read the same qualitative rationale and come to different conclusions.
A regulator or notified body wants to see that your conclusion “the benefits outweigh the risks” isn’t just a platitude, but is backed by a structured analysis and data-driven reasoning.
As one industry blog bluntly put it, qualitative arguments are inherently subjective – an issue best addressed by adding a quantitative approach.
It’s not just European regulators pushing this. FDA and other agencies worldwide also emphasize benefit-risk assessment in decision-making.
Many guidance documents now call for more transparency in how benefits and risks are evaluated, rather than a one-line statement.
In practice, a convincing benefit-risk analysis needs to lay out evidence that the device’s positive outcomes for patients meaningfully outweigh the potential harms. In the sections below, we’ll discuss how to gather that evidence and present it clearly.
Qualitative vs. Quantitative: Two Approaches to Benefit-Risk
In the real world, most medical device benefit-risk assessments historically have been qualitative or descriptive. Regulatory bodies and companies often provide user-friendly templates or worksheets to guide a qualitative discussion.
Typically, these involve listing the device’s known benefits, listing the known risks, and then writing a reasoned narrative justifying that the benefits outweigh the risks. Such an approach can work, especially for simpler devices or well-understood situations, and it’s certainly better than nothing. It allows teams to consider expert opinions and clinical context.
But it has clear limits: a purely narrative argument is still subjective and can be hard to audit for consistency. If two reviewers would likely come to different conclusions from your write-up, or if a regulator remains unconvinced, a qualitative rationale alone may fall flat.
That’s where quantitative methods come in. By incorporating some numbers or structured comparison criteria, we can introduce more objectivity and transparency into the benefit-risk conclusion.
The goal isn’t to eliminate expert judgment, but to support it with a clear framework that shows how you weighed the factors.
There is a growing recognition that we need better quantitative tools for device benefit-risk assessment.
A recent expert review noted that no current official guidance describes a quantitative method for devices, and worldwide regulators still rely mostly on descriptive approaches (pubmed.ncbi.nlm.nih.gov).
However, the same review highlighted that methodologies like Multi-Criteria Decision Analysis (MCDA), widely used for drug benefit-risk decisions, could be adapted to medical devices to bring more rigor. (We’ll discuss MCDA a bit later.)
The key takeaway is that it is possible, and often helpful, to quantify benefits and risks, even if regulations don’t spell out exactly how.
In my experience, the best results come from blending qualitative insight with quantitative evidence. The qualitative part ensures we consider context, clinical nuances, and stakeholder perspectives that numbers alone might miss. The quantitative part forces us to pin down specifics – how much, how often, how severe – and makes our argument more logical and reviewable. By combining the two, you create a narrative that is grounded in data. So, how do we actually do that? It starts with defining benefits the right way.
Defining Benefits in Measurable Terms
The first step is to clearly define the clinical benefits of your device, and to do so in the patient’s terms.
Ask yourself:
What positive outcomes does this device produce for the patient or clinician?
Is it pain relief?
Faster healing?
Fewer surgical complications?
Improved diagnostic accuracy?
Those are the kinds of benefits that matter in healthcare, and importantly, they can be measured.
In regulatory language, you want “meaningful, measurable, patient-relevant clinical outcomes” that demonstrate how the device improves health.
Some examples of measurable benefit outcomes:
Wound care device: Faster wound healing, measured by days to complete healing or percentage reduction in wound size over a certain time.
Diagnostic software: More accurate disease detection, measured by clinical endpoints like sensitivity, specificity, or accuracy in identifying a condition.
Surgical tool: Reduced complications or procedure time, measured by complication rates (% of patients with a complication) or minutes saved in surgery.
Therapeutic device: Symptom improvement, measured by a validated pain score or functional scale improvement from baseline.
It’s important to distinguish a device’s technical performance from its clinical benefit.
A device’s specs (e.g., an imaging system’s resolution, a drill’s battery life, a sensor’s precision) only matter insofar as they lead to a clinical improvement.
Regulators expect us to make that link explicit.
For example, if you claim a benefit of “better diagnostic images,” you need to tie that to a clinical outcome like improved diagnostic accuracy or earlier diagnosis that benefits the patient.
In our clinical evaluations, we actually maintain a separate list of Clinical Benefits (distinct from mere performance specs) , focusing on outcomes that improve patient health or healthcare delivery.
Each claimed benefit should correspond to a specific metric or endpoint.
In fact, EU guidance (MEDDEV 2.7/1 rev.4) insists that every clinical benefit must be associated with specific parameters that allow its evaluation (e.g. an improvement in a particular score or rate).
In short, if you list a benefit, you should be prepared to answer: How would we know, and how would we measure, that this benefit is actually being realized?
By nailing down what we mean by each benefit and how we’ll measure it, we set the stage for a more quantitative comparison later.
This exercise also forces a reality check: if you struggle to find any measurable positive outcome, that’s a red flag that the “benefit” might be trivial or unproven. On the other hand, if you have solid clinical endpoints (e.g., % of patients with symptom relief beyond a threshold, improvement on a functional scale, reduction in disease events, etc.), you’re in a good position to quantify the benefit. Remember, the MDR definition of benefit demands meaningful improvement in health outcomes.
Keep that front and center when defining what your device brings to the table.
From Outcomes to Numbers: Quantifying Benefit vs. Risk
Once we have specific outcome measures for benefits (and similarly well-defined outcomes for risks), we can start putting numbers on both sides of the equation.
Four key questions form the backbone of a quantitative benefit-risk analysis:
How big is the benefit? – i.e., what is the magnitude of improvement? (For example, does the device reduce pain by a little or a lot? Does it improve the cure rate by 5% or 50%? Increase diagnostic sensitivity from 80% to 95%?)
How severe is the risk? – i.e., what is the severity of each potential harm? (Are we talking about minor skin irritation, or a life-threatening complication?)
How often do benefits occur? – the frequency or likelihood of patients experiencing the benefit. (Do nearly all patients see this benefit, or just a subset? For instance, 80% of users have improved outcomes, or maybe only 20% respond, but with a large improvement.)
How often do risks occur? – the probability or rate of each harm. (Is an adverse event extremely rare, say 0.1% of patients, or relatively common, like 10%? Do events occur once per 1,000 uses or 1 per 10 uses?)
These factors, magnitude, and frequency for both benefits and risks, allow us to quantitatively compare the two sides. As MEDDEV guidance elegantly states: “A large benefit, even if experienced by a small population, may be significant enough to outweigh risks, whereas a small benefit may not, unless experienced by a large population.”
In other words, a rare but high-magnitude benefit might justify accepting some risk, whereas a very minor benefit would only be worthwhile if almost everyone experiences it (or if the risks are virtually nil).
This principle encourages us to weigh both how significant an effect is and how common it is.
A straightforward way to bring these elements together is a simple quantitative scoring or ratio method.
For each identified benefit, estimate:
Benefit Frequency : what proportion of patients (or users) realize this benefit.
Benefit Magnitude: how large or important the benefit is for an individual patient.
For each identified risk, estimate:
Risk Frequency: the incidence rate of that harm (how often it occurs per patient or procedure).
Risk Severity: how serious the harm is when it happens.
With these, you can calculate a Benefit Value and a Risk Value for each outcome, and then compare them:
Benefit Value = Benefit Frequency × Benefit Magnitude (on some consistent scale)
Risk Value = Risk Frequency × Risk Severity
Benefit-Risk Ratio = Benefit Value / Risk Value
This is essentially a quantified benefit-risk ratio for that particular benefit-risk pair.
For example, suppose a device significantly improves a certain outcome in 70% of patients (Benefit Frequency = 0.7, Benefit Magnitude = high).
Meanwhile, it has a moderate adverse event in 5% of patients (Risk Frequency = 0.05, Risk Severity = moderate).
We might assign a numeric score for “high” benefit (say 4 on a 5-point scale) and “moderate” severity (say 2 on a 5-point scale).
Then Benefit Value ≈ 0.7×4 = 2.8, Risk Value ≈ 0.05×2 = 0.1, giving a Benefit-Risk ratio of about 28.
In basic terms, any ratio >1.0 means the benefit value outweighs the risk value (favorable balance), whereas a ratio <1.0 would signal that risk might overshadow benefit.
In the example above, a ratio of 28 is very favorable; the benefit vastly exceeds the risk on our scale.
If we got a ratio of say 0.8, that would be a red flag that the risk might be too high relative to the benefit.
A few important notes when using this kind of approach:
Keep outcomes separate: Don’t collapse all benefits into one aggregate number or all risks into one overall risk value. If you do that, you lose crucial detail. As one guidance explains, a single composite value could allow a serious risk to be “hidden” by numerous trivial benefits or vice versa. It’s better to present a set of benefit-risk evaluations for each major benefit (paired with its most relevant risk or risks) rather than one combined ratio for the entire device. This way, you can transparently discuss each aspect of the benefit-risk profile. In fact, the experimental standard XP S99-223 (France) takes this approach: it calls for identifying “key benefits” and “key risks” that significantly affect the overall balance, and evaluating the benefit-risk ratio for each intended use scenario.
Set criteria in advance: Decide what you will consider an “acceptable” outcome of the analysis before you crunch the numbers. For instance, you might set the rule that any significant benefit should have a Benefit-Risk ratio > 1 when compared to its associated risks, or even require >2 for high-risk device scenarios. Establishing acceptance thresholds ahead of time prevents any temptation to retroactively adjust your interpretation to get a favorable result. Regulators want to see that your analysis is unbiased and rigorous, so define your pass/fail criteria up front. (XP S99-223 likewise expects manufacturers to determine acceptability criteria for benefit/risk in their policy, appropriate to each device’s context.)
Use clinically meaningful scales: If you are assigning numeric scores for magnitude or severity, anchor them in clinical reality. For example, you might define a “5” on benefit magnitude as a life-saving or life-altering improvement, whereas a “1” might be a mild symptomatic relief. Similarly, severity could be tied to standard harm classifications (like “mild inconvenience” vs “permanent impairment” vs “death”). By defining these scales, anyone reviewing your work can understand what a given score means. This improves transparency. Sometimes you’ll use direct measures (e.g., years of life saved, mm Hg blood pressure reduction) instead of arbitrary 1–5 scales – that’s even better when possible. The key is to ensure that whatever measure of magnitude/severity you use actually correlates to something that matters clinically (a 5 mm<sup>2</sup> wound size reduction might be far more impactful than a 5 mL blood loss, even though “5” of something in each case: context is everything.
Compare to the state of the art: Both EU MDR and good practice principles say that you should consider your device’s benefit-risk profile in light of other available options. In your analysis, it strengthens the case to compare your device’s benefits and risks to those of the current standard of care or a leading competitor. Ideally, your device should show an equal or better benefit-risk ratio than the existing alternative (or address an unmet need that justifies its risks). You can actually calculate the same kind of ratio for the comparator using published data: for example, if an established therapy has a Benefit-Risk ratio of ~2:1 by your method, and your new device comes out at ~3:1, that’s a solid quantitative argument that you offer an improvement. If your device’s ratio is lower than the standard’s, you will need a compelling justification (perhaps you’re targeting patients who have no other options, so some extra risk might be acceptable given the lack of alternatives). Regulators expect a discussion of the device’s benefits and risks relative to current practice as part of the overall justification.
By quantifying each major benefit and risk, we create a more objective backbone for the classic line “the benefits outweigh the risks.” It transforms that statement from a subjective claim into something you can demonstrate. This doesn’t mean the decision to approve or use a device becomes purely mathematical – far from it. Human judgment is still needed to interpret the numbers, weigh unquantifiable factors, and decide what is acceptable. But the use of data and structured comparisons makes the reasoning much clearer and more credible. In my experience, regulatory reviewers appreciate seeing this kind of analysis. It shows you’ve done your homework and aren’t just hand-waving with optimistic words.
Using NNT and NNH for a Quick Reality Check
One particularly powerful and intuitive set of metrics to incorporate into benefit-risk discussions is Number Needed to Treat (NNT) and Number Needed to Harm (NNH). These metrics, borrowed from evidence-based medicine, boil down the benefit and harm into easy-to-grasp figures: how many patients (on average) need to receive the intervention for one patient to benefit (NNT), versus how many need to be exposed for one patient to experience a harm (NNH).
NNT/NNH are essentially the inverse of absolute risk reduction or increase. They are widely used in clinical literature to communicate the trade-offs of therapies.
For example, suppose a new cardiac device prevents stroke in 5% of high-risk patients who use it (whereas without the device, those patients wouldn’t have that benefit).
That’s an absolute benefit for 5 out of 100 patients, so the NNT = 100/5 = 20.
In other words, treating 20 patients with the device prevents one stroke.
Now suppose that same device has a complication rate causing serious bleeding in 1% of patients (an absolute harm for 1 out of 100, so NNH = 100/1 = 100).
That means, on average, one patient is harmed for every 100 patients treated.
By comparing NNT and NNH, we see that for every 1 serious harm, the device prevents about 5 strokes in this scenario.
This is a strong net benefit in a life-threatening context (one harm per 5 lives saved).
If instead the numbers were closer, say NNT of 20 and NNH of 25, the balance would be questionable, and if NNH was smaller than NNT (meaning harms more frequent than benefits), that would be a red flag.
NNT/NNH give a tangible sense of the benefit-risk balance in clinical terms. They can be communicated to clinicians and even patients relatively easily (“one extra cure for every 10 treated” or “1 in 500 users experiences a severe side effect”).
Regulators and health technology assessors often consider these metrics in their evaluations. However, a few cautions:
Match the outcomes: NNT and NNH refer to specific outcomes, so you should compare an NNT and NNH that refer to outcomes of similar importance. For instance, it can make sense to compare the NNT for preventing one hospitalization vs the NNH for causing one hospitalization (because both outcomes are similar in kind). But comparing an NNT for mild symptom relief to an NNH for a life-threatening event would be misleading – those outcomes are not equivalent. If the benefits and harms differ greatly in nature, a simple ratio or comparison of NNT vs NNH needs additional weighting for importance.
Context matters: NNT/NNH are influenced by the baseline risk in the population and the timeframe considered. A device might have different NNT in different patient subgroups. So, use NNT/NNH that are relevant to your device’s indicated population and usage duration.
Despite these nuances, introducing NNT and NNH into your benefit-risk analysis is often a powerful way to make the discussion concrete. It forces clarity on absolute effect sizes and can highlight if a device’s benefit is more than just theoretical. In fact, early efforts in quantitative benefit-risk (for drugs) proposed NNT and NNH as core measures for benefits and risks, sometimes adjusting them by patient preference weights to create a “relative value adjusted NNT”. In device assessments, we can similarly use NNT/NNH as part of the toolkit to sanity-check whether the benefits meaningfully outweigh the harms in practical terms.
Other Quantitative Tools and Considerations
The simple scoring-and-ratio method above is one convenient approach, but it’s not the only game in town. For more complex benefit-risk decisions, especially if a device has multiple types of benefits and risks affecting different stakeholders, you might consider more sophisticated tools. One of the most discussed topics in recent years is Multi-Criteria Decision Analysis (MCDA). MCDA provides a structured framework to evaluate multiple benefits and risks together by assigning each outcome a weight (importance) and a score, then computing an overall benefit-risk score. It’s been used for years in pharmaceutical benefit-risk assessments to help drug approval committees make tough decisions.
Adapting MCDA to medical devices comes with some twists. A 2023 expert review (Su et al. in Expert Review of Medical Devices) noted that while regulators increasingly emphasize benefit-risk assessment, no guidance to date gives a quantitative method for devices, and most rely on qualitative worksheets (pubmed.ncbi.nlm.nih.gov.)
The authors suggest that MCDA could be one of the most useful quantitative methods for device benefit-risk, just as it has been for drugs (pubmed.ncbi.nlm.nih.gov).
They also recommend optimizing MCDA for the unique characteristics of devices by doing the following:
Use data from the current state of the art as a baseline (since devices are often compared to an existing standard of care, it’s important to bring that comparator data in).
Incorporate real-world evidence, such as post-market surveillance data and published literature, not just pre-market clinical trial data. Devices often evolve and have iterative improvements, so post-market data can be very informative.
Assign outcome weights thoughtfully, considering the type, magnitude/severity, and duration of each benefit and risk. For example, a temporary side effect might be weighted less than a permanent disability; a one-time cure might be weighted more than a symptom that requires ongoing management.
Include physician and patient opinions in the weighting process. Different stakeholders may value outcomes differently – for instance, patients might value quality of life improvements more than clinicians realize, or vice versa. MCDA allows for integrating these preference weights systematically.
In essence, MCDA lets you create a multi-dimensional benefit-risk model that can be adapted to the device’s context. Practically, implementing MCDA can be as involved as setting up a decision-analytic model or as simple as a weighted pros-and-cons table, depending on the level of rigor needed. If you go this route, it’s wise to follow established good practices (ISPOR has guidance on using MCDA in healthcare decisions). The outcome of an MCDA is often an overall score or ranking indicating whether the benefits outweigh the risks, given the weights applied.
Another consideration is the use of composite metrics like Net Health Benefit or Quality-Adjusted Life Years (QALYs). These are common in health economics and health technology assessment: for example, combining multiple effects into a single scale of “health utility”. For devices that have a significant impact on survival or quality of life, you could theoretically quantify benefit in terms of QALYs gained and risk in QALYs lost (due to side effects), yielding a net value. This is essentially a form of MCDA where the common currency is a utility value. However, doing a full QALY-based analysis can be data-intensive and may not be expected unless you are preparing a reimbursement dossier or a very high-stakes regulatory submission.
Even if you don’t deploy formal MCDA or QALY models, you can make your qualitative analysis more structured. One way is to use a checklist of factors to discuss (with evidence for each). For instance, the FDA’s benefit-risk framework for premarket approvals suggests discussing:
The severity of the condition the device addresses (more severe disease might justify higher risk).
Unmet medical needs (if patients have no alternatives, you might tolerate more risk).
The type and magnitude of the device’s benefits (what health aspects improve, and by how much).
The probability of patients experiencing benefits (e.g. responder rate).
The severity and probability of risks.
Patient tolerance for risk and perspective on benefit (are patients generally willing to accept the risks for the promised benefit?).
Risk mitigations and how uncertainty is handled (confidence in the data).
Ensuring your benefit-risk narrative covers these elements in a systematic way, ideally with quantitative data for each, can transform a fuzzy discussion into a convincing evidence-based argument. It also demonstrates that you’ve thought about the issue from multiple angles, which is exactly what regulators want to see. In fact, involving patients in the process is increasingly encouraged. The French standard XP S99-223 explicitly calls for consideration of patient opinions as part of benefit-risk management, including their perception of benefit importance, risk severity, and acceptable trade-offs. Techniques like patient surveys or preference studies can be used to gauge what level of risk patients would tolerate for a given benefit. Incorporating such insights not only strengthens your analysis but aligns with the regulatory trend toward patient-centric evaluation.
The bottom line is that there are many tools available, from simple arithmetic ratios to elaborate decision models, to bring more objectivity into benefit-risk assessment. The choice of tool should be proportionate to the complexity and risk of your device. A Class I device with minor effects might only need a well-reasoned qualitative discussion (but still structured and with some data). A novel life-sustaining implant might warrant a full quantitative model with multiple weighted criteria. Whichever approach you choose, document it clearly and justify why it’s appropriate. The worst outcome is to have a complex analysis that no one can understand; the best outcome is an analysis that is as simple as possible, but not simpler than necessary to capture the true balance of benefits and risks.
Case Example: Quantifying Benefit-Risk for an AI Diagnostic Device
To illustrate how one might blend qualitative and quantitative approaches, let’s walk through a hypothetical example. Imagine we are developing an AI-based software medical device, call it SmartScanAI, designed to assist in detecting early-stage lung cancer from chest X-rays. This kind of tool would be used by clinicians in general practice or radiology to flag suspicious nodules that might indicate cancer, prompting further evaluation. We need to assess the benefits against the risks:
1. Define the Benefits and Risks (Qualitatively and in Outcomes):
From a patient’s perspective, the key benefit of SmartScanAI is earlier and improved detection of lung cancer, which can lead to earlier treatment and better survival.
Specifically, we’d define a measurable benefit outcome such as “increase in the detection rate of actionable lung nodules/cancers”. For instance, if traditionally 50% of lung cancers in their early stage are detected on X-ray by human readers, does our AI bump that to 70%?
Another potential benefit might be diagnostic efficiency; maybe the AI helps radiologists read images faster (though that’s more of a workflow improvement; we’ll focus on clinical outcome as the primary benefit).
On the risk side, SmartScanAI doesn’t directly touch the patient, so it has no physical side effects in and of itself. However, it can produce indirect risks through erroneous outputs:
False positives: The AI might incorrectly flag benign spots as suspicious. This could lead to unnecessary follow-up tests like CT scans or even invasive biopsies, which carry their own risks (radiation exposure, biopsy complications, patient anxiety).
False negatives: The AI might miss a cancer (especially if the human reader becomes over-reliant on the AI). This could delay diagnosis, potentially harming the patient by missing the window for early treatment. (In practice, the radiologist is still responsible, but over-reliance is a human factor risk to consider.)
For our analysis, let’s focus on the quantifiable effects of false positives leading to unnecessary procedures, since false negatives are harder to quantify (their harm is essentially “benefit not achieved,” which we can address by looking at residual missed cancers).
2. Gather Data and Estimates:
Suppose in a clinical study or pilot trial of SmartScanAI involving 1,000 patients, the results were:
Benefit data: With AI assistance, 50 early-stage lung cancers were detected out of 1,000 patients screened (5%). Without AI (historical control or parallel group), only 40 would have been detected (4%). So the AI led to 10 additional early cancer detections per 1,000 patients, an absolute improvement of 1% (relative increase of 25% over the baseline 4%). These 10 patients, by getting earlier treatment, have significantly better prognoses than if their cancers were caught later. For simplicity, let’s assume early detection in these cases translates into lives saved or significantly extended (early-stage lung cancer 5-year survival is much higher than late-stage). This is our primary benefit: 10 extra patients out of 1000 benefit by potentially life-saving early diagnosis. We can also express that as NNT = 1000/10 = 100, meaning 100 patients need to be screened with SmartScanAI for one additional early cancer to be found (one life potentially saved). An NNT of 100 might sound high, but for a screening context and a serious disease, it’s actually quite reasonable, consider that many population screening programs for cancer have NNTs in that ballpark or higher when the disease is rare. The magnitude of the benefit (for those 10 patients) is very high: a life-threatening illness caught early. So we’d rate the benefit magnitude as “major” for those who benefit, and the frequency of benefit as 1% (for the population overall).
Risk data: The AI flagged a lot of images as “suspicious” – say 200 out of 1000 patients had some finding flagged. Besides the 50 actual cancers among those (true positives), there were 150 false positives. Not all false positives will lead to harm; typically, a suspicious X-ray leads to a CT scan. So 150 patients got follow-up CTs that turned out normal. CT scans have radiation exposure (small risk of inducing cancer decades later) and cost, but the more immediate potential harm is if a fraction of those patients undergo an invasive procedure like a biopsy. Let’s assume, of the 150 false positives, about 10 ended up getting a needle biopsy of a lung nodule “just to be sure.” Biopsy is an invasive procedure with risks like pneumothorax (collapsed lung) or bleeding. Suppose out of those 10 unnecessary biopsies, 1 patient had a significant complication (e.g. a pneumothorax requiring hospitalization). That is 1 serious harm per 1000 patients screened due to the AI’s false positives. In terms of NNH, that’s NNH = 1000 since one harm occurred for 1000 uses. We could also include minor harms (maybe some patients had anxiety or smaller pneumothoraces that didn’t need intervention, etc., but let’s focus on a significant harm). The risk frequency of a serious complication here is 0.1% (1/1000), and the severity of that harm we’d rate as moderate or serious (let’s say it’s serious but not life-threatening, so on a severity scale maybe 3 out of 5).
Now we have quantitative pieces:
Benefit: 1% of patients get a major life-saving benefit (Benefit Frequency = 0.01, Benefit Magnitude = very high).
Risk: 0.1% of patients get a serious harm (Risk Frequency = 0.001, Risk Severity = moderately high).
3. Quantify and Compare:
Using the Benefit = Frequency × Magnitude approach, we might assign a numeric magnitude score for convenience. If we say “saving a life/major outcome” is a 5 on a 1–5 scale, then Benefit Value = 0.01 × 5 = 0.05 (per patient screened). For the harm, if a serious complication (hospitalization but recoverable) is maybe a 3 on severity, then Risk Value = 0.001 × 3 = 0.003. Now compare: the Benefit-Risk ratio = 0.05 / 0.003 ≈ 16.7. This ratio >> 1, indicating a strongly favorable balance, the benefit (in aggregate probability-weighted terms) is about 17 times greater than the risk by our scoring. If we prefer to think in NNT/NNH terms: NNT = 100, NNH = 1000. For every 1 harm caused, 10 patients benefit (or equivalently, one harm per 10 benefits). This is another way to see that the scales tilt toward benefit.
Of course, we must interpret what this means. In this scenario, SmartScanAI’s benefits likely outweigh its risks. The benefit is life-saving for a few patients out of 1000, while the harm is a serious but not life-threatening complication in one patient (which, with proper care, the patient recovered from). Many would argue that preventing ten cases of advanced cancer is well worth one iatrogenic complication. We would document it like: “SmartScanAI improved early lung cancer detection by 25% (catching an additional 1% of the screened population in early stage)【(hypothetical data)】. This translates to an NNT of ~100. In exchange, it led to an unnecessary invasive procedure in some patients, with an incidence of serious complication in 0.1% (NNH ~1000). Given that the harm (treatable pneumothorax) is considerably less severe than the benefit (potentially life saved or prolonged), and is much less frequent, the overall benefit-risk is favorable (quantitatively, the benefit-risk ratio ≫1 in this population).”
We would also compare to the standard of care without AI. Without AI, by definition, those 10 extra cancers would not be caught early, so the standard of care’s benefit-risk profile would be poorer in terms of benefit (fewer lives saved).
The standard of care has essentially no “AI-induced harms,” but it has the harm of missed cancers (which is effectively the inverse of the AI benefit). So we could say: “Compared to not using the AI, using SmartScanAI results in 10 additional early cancer detections per 1000 (major benefit), at the cost of 1 additional procedure-related complication per 1000 (moderate harm).
No current alternative method provides such an improvement in early detection in this setting, so SmartScanAI offers a net positive clinical impact.” This context is important; if there were another software that catches 15 extra cancers with only 1 complication, that would be even better, but assuming our AI is the state of the art improvement, it justifies its adoption.
4. Document Rationale and Mitigations:
In writing up the benefit-risk analysis, we wouldn’t just drop the numbers and call it a day. We would explain assumptions and consider uncertainties. For instance:
Acknowledge that the study sample was 1000 patients; as we collect more data post-market, these rates might change (we’ll update the analysis accordingly).
Perhaps the false-positive rate can be improved with algorithm tweaks or user training, further reducing the NNH (this would be part of risk control considerations).
We should mention what is an acceptable threshold in this context. Maybe we decide that any life-saving benefit that helps at least 1 per 200 patients (NNT ≤ 200) is worth any harm that is less frequent than 1 per 50 patients (NNH ≥ 50) in a cancer screening scenario. In our case, we easily meet that criterion (NNT 100, NNH 1000). Stating such criteria shows how we judged acceptability.
We also discuss patient preference: many patients, if asked, might accept a 0.1% chance of a complication in exchange for a 1% chance of early cancer detection that could save their life. That perspective supports our conclusion. In fact, if patient focus groups indicated they are comfortable with the trade-off, we would mention that as further justification (aligned with considering patient opinion per XP S99-223 and regulatory expectations).
By quantifying the benefit and risk in this example, we turned a potentially vague claim (“the AI improves diagnosis with minimal risk”) into a concrete, reviewable statement: “The AI increases early cancer detection by 10 per 1000 (NNT=100), while causing serious complications in 1 per 1000 (NNH=1000). Given the life-saving nature of early detection, this tenfold difference in favor of benefit, and the relative severity of outcomes, we conclude the benefits outweigh the risks.” This kind of analysis would make it much easier for an independent reviewer to follow and agree with our logic, compared to a generic assurance. It also helps our own team make informed decisions (for instance, if the numbers came out differently, say NNT=100 and NNH=50 – we might reconsider launching the product or going back to improve the algorithm).
5. Continuous Improvement:
One final note: the analysis doesn’t end here. Once SmartScanAI is on the market, we would collect real-world data: How many cancers are we actually catching? Are unexpected issues popping up? Perhaps the complication rate was higher or lower in broader use. We should be prepared to update the benefit-risk assessment with post-market surveillance findings.
Benefit-risk is not a one-time checkbox for the regulatory file; it’s meant to be a living assessment that gets refined as evidence accumulates (this is a requirement in many jurisdictions as part of ongoing risk management and Post-Market Clinical Follow-up). In our example, if a rare but serious risk (like a clinical decision error attributable to the AI) emerged in the field, we’d factor that in and see if the benefit-risk still holds or if mitigations are needed.
Making Benefit-Risk Assessment Credible and Transparent: Key Takeaways
Whether you’re dealing with a simple gadget or a high-risk implant, a credible benefit-risk assessment comes down to clarity, evidence, and balance. Here are some key takeaways to keep your analysis on solid ground:
Start with the patient in mind: Define benefits in terms of outcomes that matter to patients (or clinicians). Ask “What tangible improvement does this device provide to health or quality of care?” and focus on that. If your listed “benefits” are all technical specs (speed, resolution, etc.), translate them into patient-centric outcomes (faster diagnosis, more accurate treatment, etc.). Regulators will focus on whether your device actually improves patient health, rather than just performing as engineered.
Make it measurable: For each benefit you claim, identify how you can measure it and gather data to support it. Numbers speak louder than adjectives. Instead of saying “improves healing time,” say “reduces median wound healing time from 6 weeks to 4 weeks” – and back it up with a study or literature. Every benefit should ideally have a statistic or clear metric attached. Similarly, quantify risk rates from clinical data or published benchmarks (e.g. “2% risk of device-related infection”). Measurable outcomes make your benefit-risk argument concrete.
Align benefits with risks (compare apples to apples): Don’t discuss benefits and risks in completely separate silos. It’s most illuminating to pair each major benefit with the risk(s) most relevant to it. For example, if your device’s benefit is avoiding a certain complication of a disease, then discuss any device-related complications in that context. If your device provides faster results, weigh that against any risks introduced by that speed or by false results. By aligning them, you ensure you’re doing a fair comparison. This also helps in creating those benefit-risk pairs for quantification. Presenting a side-by-side comparison (e.g. in a table of Benefit vs. Corresponding Risk, with frequencies and severities) can be very effective.
Use a structured method (and quantify when possible): Bring structure to your analysis. This could be a formal quantitative method (like calculating benefit/risk ratios for key outcomes, or performing an MCDA), or a qualitative framework that covers all the important factors systematically. Wherever feasible, introduce quantitative estimates, even if they are rough. Consider using a simple scoring scheme or metrics like NNT/NNH to support your points. A small table of “Benefit X: affects Y% of patients, magnitude = high (e.g. 30% symptom improvement); Risk A: occurs in Z% of patients, severity = low” immediately adds objectivity. Numbers force you to be specific and help reviewers see your rationale clearly.
Justify your ratings and thresholds: If you assign any subjective ratings (e.g. calling a benefit “major” or a harm “negligible”), explain what those terms mean. For instance, you might define “major clinical benefit” as something that significantly improves patient survival or functioning (with an example), whereas “minor benefit” might be a modest symptom improvement. Likewise, define levels of harm severity (using, say, the Common Terminology Criteria for Adverse Events, or similar). By providing these definitions, you show that your scoring isn’t arbitrary. Also state any threshold you use (e.g. why you consider a B/R ratio >1.5 as acceptable in your case, or why a certain risk rate is considered negligible based on clinical guidelines). This transparency makes your conclusions much more robust and credible.
Compare to alternatives (state of the art): Always contextualize the benefit-risk profile by comparing it to the current standard of care or competitor devices. Is your device safer, more effective, or addressing a need unmet by existing solutions? If your benefit-risk is clearly better than the alternative (e.g. lower complication rate for the same benefit, or greater benefit for similar risk), highlight that – it’s a strong argument for why your device should be on the market. If it’s roughly on par, that’s still useful to mention (it means the device is at least as good as what’s already accepted). And if it’s worse in some aspect, be honest and explain why it might still be justified (perhaps it’s for patients who can’t use the other device, etc.). Regulators explicitly expect a “state of the art” comparison in your benefit-risk analysis, as per MDR and guidance.
Keep it honest and balanced: A credible analysis doesn’t try to sweep negatives under the rug. Acknowledge the device’s limitations and any uncertainties in the data. If a certain risk is serious but rare, don’t ignore it – explain how rare it is, how you’re mitigating it, and why the benefits still outweigh it. Conversely, avoid overselling a benefit; if it’s small, say so, but also note why even a small benefit could matter (e.g. if the condition is very serious or if cumulative impact is large across a population). By showing that you’ve weighed the pros and cons even-handedly, you build trust. Reviewers are adept at spotting one-sided arguments, so preempt their concerns by addressing them yourself. Think of the “silent skeptic in the room” and answer their questions in your report.
Incorporate patient and clinician input: Whenever possible, bring real-world perspectives into the analysis. This could be qualitative (e.g. quotes or survey data on what patients value or fear) or quantitative (preference study results). For example, if patients say they would accept a 10% risk of a minor surgery for a 50% chance at pain relief, and your device’s risk/benefit is within that range, that’s compelling evidence. Patient-centric assessment is encouraged by standards like XP S99-223 and can make your benefit-risk evaluation more persuasive by aligning it with actual user values.
Iterate with new data (benefit-risk is not static): Treat the benefit-risk assessment as a living evaluation. During development, revisit it whenever significant new data comes in (e.g. results from a new clinical study). After market launch, continue to monitor. Regulations like ISO 14971:2019 and EU MDR require that we update our conclusions with post-market information. Maybe your device in real use is performing even better than expected (great, update the benefit side with that!). Or perhaps a previously unknown infrequent risk surfaces – include it and see if the balance remains acceptable, and take action if not. Regularly scheduled benefit-risk reviews (e.g. part of annual Post-Market Surveillance reports or Periodic Safety Update Reports) ensure that “the benefits outweigh the risks” remains true throughout the device’s life cycle. This proactive approach also prepares you well for any regulator questions or field safety corrective actions, since you’ll have a quantified rationale ready if tough decisions (like a recall or an indication restriction) need to be made based on benefit-risk.
In conclusion,
making your benefit-risk analysis quantifiable and transparent is a mindset that can improve decision-making at every stage of device development.
By clearly delineating how your device helps patients and what downsides come with it, and backing those points with data, you not only satisfy regulators but also gain insights that could guide product improvements and risk mitigations. In my consulting work, teams often discover through this process that they could enhance a certain benefit or reduce a particular risk, thereby shifting the balance further in favor of patients.
So, next time you find yourself writing “In conclusion, the benefits outweigh the risks because of X, Y, Z...,” challenge yourself to put a number or piece of evidence next to each X, Y, Z.
Even if you start with just one measurable benefit and one quantifiable risk, you’ll see the power of this approach. It turns a potentially skeptical “prove it” from a reviewer into a well-supported discussion. And who knows – you might even impress that next regulatory reviewer with a benefit-risk assessment that doesn’t require a crystal ball to believe in it.
Let’s move benefit-risk from being a dull, obligatory section of the paperwork into a true, quantifiable demonstration of our device’s value to patients.
After all, if we can’t clearly articulate why the benefits outweigh the risks, we can hardly expect regulators or customers to be convinced. By applying the principles above, patient focus, measurable outcomes, structured comparisons, and continual updates, we can ensure our benefit-risk conclusions are not just hopeful statements, but solid, reviewable conclusions that stand up to scrutiny.
This whole methodology traces back to 2022, when I received a non-conformity from a Notified Body regarding my benefit-risk evaluation. That triggered a deep search, which led me to the French experimental standard XP S99-223 and a post from RQM+, both of which inspired me and helped reshape my entire approach. So, thanks to all of them. 🙌
✌️ Peace,
Hatem
Your Clinical Evaluation Expert & Partner


Thank you for an excellent article! Your summary of the regulatory requirements for assessing benefit-risk and of state of the art were both spot on. And, to top it off, you offered a significant proposal on how to improve assessments of benefit-risk. I particularly liked how symmetrically you treated benefit and risk, as this is a necessary first step to comparing their size.
I also want to make you aware of a novel approach to assessing benefit-risk that not only extends each of the positive suggestions you made for assessing benefit-risk but reaches a completely general solution. This solution is both dramatically more objective than the SOTA and provides traceable logic to show why benefit exceeds risk or risk exceeds benefit.
Like you, this method starts by defining benefit in a clinically relevant way. Specifically, the method starts from FDA's definition of a 'patient' as someone with a specific health concern, whether or not they are currently being treated for that concern. If we describe the health concern in terms of risks to the patient's health, then the 'benefit' of a treatment' becomes the difference in the risks to the patient's health before and after treatment.
Sincerely, Richard Matt
Richard.matt@aspenmedicalrisk.com
(m) (408) 483-8261
Plot twist: turns out Bethany Chung and I ghostwrite for Substack now
So there I was, scrolling through Substack, when I came across an article that gave me intense déjà vu. Every sentence. Every structure. Every carefully crafted regulatory insight. It was like looking in a mirror... written by someone else. Turns out someone copied, pasted, and made it their own. I guess Bethany and I should feel flattered? Although “flattered” is a bit strong when your white paper gets abducted, repackaged, and published with all the originality of a knock-off handbag.
To be clear: This wasn’t “inspired by” or a “summary of”. This was a full-on Ctrl+C / Ctrl+V without Ctrl+Conscience.
Here’s the original piece we wrote back in 2022: https://www.rqmplus.com/blog/a-quantitative-approach-to-benefit-risk-determination-rqm/
I mean, I get it. Regulatory writing is hard. Original thought takes time. But come on, at least change the font? At RQM+, we stand for scientific rigor and integrity. So here’s a gentle reminder: Plagiarism isn’t a literature review. Borrowing without credit is pure subterfuge and not thought leadership.
If anyone needs help writing actual content, we @RQM+ do consulting. Ethically.