Open Question

What are the hallucination rate, calibration, and clinically-relevant failure modes of vision-language/foundation models on real neuro-oncology images?

Evidence grade: WEAK — Supported only by weak or preliminary evidence.

Why it matters

These models are entering clinical view with fluent output but unquantified safety; premature trust is a patient-safety risk.

Conflicting evidence

Generalist-AI promise vs GPT-4V scoring 47.8% on image-based radiology questions; failure characterization for neuro is essentially absent.

What would resolve it

Adversarial and prospective evaluation on curated neuro-oncology cases; report calibration, abstention, and error taxonomy, not just accuracy.

Open in the interactive atlas →

About this page

Curated and maintained by Resonant Labs. Reviewed by Claims are synthesized from the cited literature and graded per the Atlas methodology (about the Atlas). Last updated 2026-07-01.

This is an educational research resource from Resonant Labs, not clinical advice. Evidence grades and controversies are summarized from the literature and may change as the field evolves.