Most hospital executives in Asia are no longer asking whether artificial intelligence belongs in clinical settings. They are asking a harder question: how do you tell a useful tool from a well-marketed one before signing a multi-year contract?
The honest answer is that the evidence base is thinner than the sales decks suggest, and the gap between “regulator-cleared” and “proven to help patients” is wide enough to fall into. Understanding that gap is now a core competency for anyone running a hospital, a clinic group or a diagnostics service.
Regulatory clearance is not evidence of patient benefit
This is the single most misunderstood point in healthcare AI procurement, and the numbers are stark.
A 2026 evidence census published in PLOS Digital Health reviewed every AI and machine-learning enabled medical device cleared or approved by the US Food and Drug Administration through 5 December 2025. Of 1,357 cleared devices, only 34 were linked to registered prospective trials, 12 had peer-reviewed publications, and just three had been evaluated against patient-centred outcomes such as mortality, morbidity or readmissions. Most of the studies that did exist used observational designs with small, homogeneous cohorts and frequently excluded vulnerable populations.
:antCitation[]{citations=”9febdc32-bbc3-452a-9f37-dc242b975a2c,72d6d138-5579-4ca3-9c0e-d00059e1c82d”}
A separate cross-sectional analysis in JAMA Network Open examined 903 FDA-approved AI-enabled devices. Clinical performance studies were reported for 505 devices, while 218 explicitly stated that no performance study had been conducted. Retrospective designs dominated; only 41 studies were prospective and 12 were randomised. Fewer than a third of clinical evaluations reported sex-subgroup information, and under a quarter reported age subgroups.
:antCitation[]{citations=”fbf27223-5c49-4c91-9c39-cbb7cd253346″}
Why does this happen? Largely because of the pathway most of these products travel. The majority reach market through a route that requires a manufacturer to demonstrate substantial equivalence to an existing device rather than to prove clinical effectiveness prospectively — which allows evidence gaps to pass down chains of predicate products.
:antCitation[]{citations=”bfd479c3-4891-47ab-ac5b-626f4a4fb497″}
None of this means AI tools do not work. It means clearance is a floor, not a verdict — and the burden of asking for outcome evidence sits with the buyer.
Asia’s regulatory picture is converging, but not uniform
Health systems across the region are moving at different speeds and with different classifications. A tool that is straightforward to deploy in one market may face a substantially heavier evidentiary burden in another.
| Market | Current position | What it means for a deploying hospital |
|---|---|---|
| Singapore | MOH and HSA published Artificial Intelligence in Healthcare Guidelines (AIHGle) 2.0 in March 2026, building on the 2021 edition | Responsibilities are explicitly split between developers, deployers and clinical users across the product lifecycle — the hospital carries defined obligations of its own |
| South Korea | Digital Medical Products Act in force since 24 January 2025, with dedicated guidance on generative AI-enabled devices | Clearer approval expectations for generative tools; a useful reference benchmark for buyers elsewhere in the region |
| China | NMPA technical review guidelines for AI medical devices; AI-based auxiliary diagnostic software generally treated as high-risk | Higher classification means a heavier evidence package, which buyers can request and read |
| Japan | PMDA has developed an adaptive framework for AI-enabled software | Post-approval change management is formally anticipated rather than improvised |
AIHGle 2.0 is intended to give practical guidance for the safe development, deployment and use of AI in healthcare, and it strengthens accountability by clarifying who is responsible for what across developers, healthcare organisations and healthcare professionals. It applies across AI broadly but targets the more complex machine-learning and deep-learning subset, where opacity and scale amplify risk. For hospital boards anywhere in Southeast Asia, it is worth reading even outside Singapore — it is the clearest articulation in the region of what a *deployer* is expected to do.
:antCitation[]{citations=”e6df2091-4cdd-42bf-bd0b-72f06a2fef99,e75590b9-a5f4-4b45-a6f5-d48542c748b9″}
The World Health Organization has taken a similar line. Its January 2024 guidance on large multi-modal models set out more than 40 recommendations spanning governments, technology companies and healthcare providers, and it addresses how health systems should evaluate, procure and oversee these tools while managing safety, bias, privacy and accountability.
:antCitation[]{citations=”a08ac73e-0974-4604-a041-b01fcb400c59,c6f30623-ceae-46bb-b8bf-23437caeb007″}
Six questions to ask before deployment
This framework is deliberately ordered. Each question is harder to answer than the last, and a vendor who stalls at question three has told you something useful.
- What exactly is the intended use, stated narrowly? “Assists radiologists” is not an intended use. The specific modality, body region, patient group and clinical decision point are.
- What evidence exists, and of what type? Ask whether performance data is retrospective or prospective, internal or external, single-site or multi-site. Retrospective single-site accuracy is a starting point, not a result.
- Was the tool validated on populations resembling ours? This is where Asian health systems have a specific interest. Disease prevalence, screening pathways, imaging equipment and population genetics differ from the cohorts where most models were trained. Published analysis has warned that tools authorised in high-income settings are routinely deployed elsewhere without contextual validation, and cited historical examples where concordance with local clinical practice fell far below initial reported performance.
- What happens when the model changes? Software updates silently; clinical performance can drift. Ask for versioning, a change-control plan and defined re-validation triggers before go-live, not after.
- Who is accountable when the output is wrong? Establish in writing where clinical responsibility sits, what the escalation path is, and how a clinician overrides the tool without friction.
- How will we measure whether it actually helped? Agree the metric before deployment. Time saved, recall rate, missed-finding rate, throughput — whatever it is, define it, baseline it, and commit to publishing the result internally even if it is unflattering.
:antCitation[]{citations=”6126a8ae-9ca8-43d7-8c7e-0bba9201a0b1″}
Four different claims — and the evidence each one needs
Much of the confusion in AI procurement comes from treating four very different statements as if they were interchangeable. They are not, and they require entirely separate evidence.
| Type of claim | Example | What actually supports it | What it does not prove |
|---|---|---|---|
| Regulatory | “Cleared as a medical device” | A regulator’s marketing authorisation | That patients do better |
| Technical | “96% accuracy on our test set” | A dataset and a metric | Performance in your hospital, on your patients |
| Operational | “Deployed across 40 sites; 2 million studies processed” | Auditable deployment and volume records | Clinical effectiveness |
| Clinical | “Reduces missed diagnoses” | Prospective, ideally multi-site, outcome studies | — |
A tool can be entirely legitimate at the first three levels and still have no evidence at the fourth. That is not necessarily a reason to reject it. It is a reason to be precise about what you are buying and what you tell your clinicians.
Where measurable milestones fit — and where they don’t
Operational scale deserves its own note, because healthcare organisations increasingly want credit for it, and increasingly deserve it.
When a hospital group completes an unusually large AI-assisted screening programme, or a healthtech company reaches a documented regional deployment milestone, that achievement is real, countable and independently verifiable. It belongs to the same family as other forms of healthcare recognition such as accreditation and awards — and like them, it answers a specific question rather than every question.
Independent record recognition sits in this space. Organisations such as Asia Record maintain a public register of verified record holders, and healthcare, wellness and medical technology companies appear among them. For a hospital group or healthtech company, becoming an Asia Record holder documents a measurable organisational milestone — participation volume, programme scale, geographic footprint, a first-of-its-kind operational achievement. That is a legitimate form of corporate record recognition in Asia, and it is verifiable in a way that most marketing language is not.
What it does not do — and what no responsible organisation should imply — is establish clinical superiority. Record certification in Asia evidences scale and verifiability, not therapeutic effectiveness. A healthcare business achievement and a clinical outcome claim are different assertions requiring different proof, and conflating them damages the credibility of both. Organisations exploring how to get an Asia Record for a documented institutional milestone can review the official Asia Record application and nomination process, which sets out eligibility, documentation requirements and review timelines. The discipline it imposes — specific, measurable, evidenced, independently reviewed — is the same discipline that should govern every AI claim a hospital accepts.
A practical pre-deployment checklist
- Written intended-use statement, narrowly scoped, signed off by the clinical lead who will use it
- Full evidence dossier requested in writing, including study designs and site counts — not a summary slide
- Explicit question on subgroup performance, including whether the validation population resembles yours
- Local shadow-mode evaluation period before any tool influences a clinical decision
- Documented change-control and re-validation policy covering vendor updates
- Named clinical accountability and a low-friction override pathway
- Pre-agreed success metrics, baselined before go-live
- Post-deployment monitoring schedule with defined triggers for suspension
- Data governance and cybersecurity review aligned to your national framework
The reasonable position
Scepticism is not the same as refusal. Some AI tools in imaging, triage and administrative workflow are producing measurable operational benefit right now, and hospitals that ignore them will fall behind. The point is that the burden of proof has quietly shifted onto buyers, and most procurement processes have not caught up.
The institutions that will handle this well over the next few years are not the fastest adopters. They are the ones that ask better questions, document what they deploy, measure what happens afterwards, and describe their achievements accurately — separating what they have built from what they have proven. That is the same standard the region’s more credible healthtech companies are already applying to themselves before scaling, and it is increasingly what sophisticated buyers, regulators and clinicians expect.
This article is intended for healthcare professionals and organisations and discusses institutional technology governance. It is not clinical guidance and does not replace professional medical or regulatory advice.