Open Question

What are the hallucination rate, calibration, and clinically-relevant failure modes of vision-language/foundation models on real neuro-oncology images?

Evidence profile

Status: question. Rule: atlas-profile-2.0.

5 linked publications; not a count of independent studies.

Missing appraisal does not mean no evidence exists. Inspect the full profile and source relationships in the interactive Atlas.

Historical grade: WEAK (retained from 22 September 2026; not a completed profile 2.0 assessment).

Why it matters

These models are entering clinical view with fluent output but unquantified safety; premature trust is a patient-safety risk.

Conflicting evidence

Generalist-AI promise vs GPT-4V scoring 47.8% on image-based radiology questions; failure characterization for neuro is essentially absent.

Proposed study direction

Adversarial and prospective evaluation on curated neuro-oncology cases; report calibration, abstention, and error taxonomy, not just accuracy.

Prepare a research brief

Gap verification

This question comes from selected literature. Confirm that the gap remains open with a current search and mentor review.

Open in the interactive atlas →

About this page

Curated and maintained by Resonant Labs. Editorial lead: Claims are synthesized from the cited literature and graded per the Atlas methodology (about the Atlas). Page generated 2026-10-02. This is not a new clinical review date.

This is an educational research resource from Resonant Labs, not clinical advice. Atlas assessments are heuristic, not probabilities or formal GRADE ratings. Findings are summarized from the literature and may change as the field evolves.