Alibaba DAMO Academy Open-Sources General-Purpose Abdominal CT AI Capable of Identifying 146 Conditions
nashnova research
DAMO Academy has open-sourced DAMO RADAR, a general-purpose medical imaging AI that reads a single abdominal contrast-enhanced CT to flag 146 conditions across 18 organs — with the paper published in *Science*, this marks the step from lab concept to deployable open-source tool.
What does this model do, and why call it "general-purpose"?
DAMO RADAR is a vision-language model — trained on paired CT scans and clinical text reports — that screens 18 abdominal organs for 146 clinical findings in one pass, including multiple malignancies.
This means → it is not a single-disease detector but a one-scan, full-spectrum screener. DAMO Academy calls it "the world's first expert-level general medical imaging model."
Across nearly 39,160 real-world exams, its average AUC — a metric where 1.0 is a perfect diagnostic score — reached 0.913. In plain terms = in roughly 9 out of 10 judgments the model correctly separates "disease present" from "disease absent," approaching top-tier radiologist performance.
How did it learn to read scans "like a doctor"?
DAMO senior algorithm researcher Zhang Jianpeng spent extended time observing radiologists at work and found the key behavior: doctors do not scan the entire abdomen at a glance — they check organ by organ.
RADAR's architecture mirrors that routine: it splits the abdomen into anatomical units, aligns each organ with its corresponding text description in the report, then compares positive and negative text prompts against the image to generate a heatmap — a color-coded risk overlay — for the radiologist.
This reflects a deliberate design trade-off. Earlier vision-language approaches took a "global view" of abdominal CT and underperformed because the abdomen is structurally complex and lesions are sparse — the model learned shortcuts, and critical abnormalities got buried. Organ-level decomposition solved that problem.
How does it compare with real doctors?
Tested against 26 radiologists from 14 hospitals, RADAR's average accuracy exceeded 23 of them and fell short of only 3 senior specialists.
On 27,267 emergency CT cases covering 17 common acute abdominal conditions, AUC still reached 0.904; retrospective validation across 8 external medical centers yielded an average AUC of roughly 0.895.
In assisted-reading trials, AI raised doctors' sensitivity — their ability to avoid missed diagnoses — by 10% and cut average reading time by 30.7%. This means → the model's biggest value is not replacing doctors but helping them work faster with fewer misses.
What do frontline doctors think — tool or crutch?
Xiao Wenbo, radiology chief at the First Affiliated Hospital of Zhejiang University, noted the department handles 4,000–5,000 CTs a day; a single abdominal CT takes a senior radiologist 20–30 minutes. "If abdominal CT is already tough at our hospital, it must be harder at smaller ones."
She also flagged a risk: junior doctors aided by AI can produce reports near senior-level quality, "but once you take the AI away, they're back to junior-level."
Her proposed workflow: doctor makes an independent call → AI cross-checks → doctor compares and reconciles — AI as a teaching aid, not a replacement. In plain terms = do the homework first, then check the answer key — don't copy it.
How does RADAR relate to DAMO's earlier single-disease models?
RADAR is not a simple merger of DAMO PANDA (pancreatic cancer), GRAPE (gastric cancer), COCA (colorectal cancer), and EAGLE (esophageal cancer).
Those single-disease models scan non-contrast CT for early signals of a specific cancer, suited to opportunistic screening in asymptomatic populations. RADAR targets contrast-enhanced CT and aims to check multiple organs and a wide range of abnormalities simultaneously. This means → the two serve different roles: one is "screen while you're getting a checkup," the other is "full workup once symptoms appear."
DAMO senior algorithm researcher Zhang Ling said RADAR's capability "has not been fully tapped and could continue to rise with scaling laws."
It's open-sourced — how far is it from real clinical use?
Code is on GitHub, the paper is in *Science*, and deployment barriers for outside institutions have dropped sharply.
Yet RADAR has only completed retrospective studies — validation on existing data. It has not entered prospective clinical trials, where it would be tested in real-time diagnostic workflows.
This reflects the central open question for general-purpose imaging AI: retrospective numbers look strong, but whether they hold up in prospective trials is the make-or-break validation milestone. In plain terms = acing a practice exam is promising, but the real test is the live one.
市场有风险,内容仅供研究参考,不构成投资建议。
