Evidence grade: WEAK — Supported only by weak or preliminary evidence.
These models are entering clinical view with fluent output but unquantified safety; premature trust is a patient-safety risk.
Generalist-AI promise vs GPT-4V scoring 47.8% on image-based radiology questions; failure characterization for neuro is essentially absent.
Adversarial and prospective evaluation on curated neuro-oncology cases; report calibration, abstention, and error taxonomy, not just accuracy.
This is an educational research resource from Resonant Labs, not clinical advice. Evidence grades and controversies are summarized from the literature and may change as the field evolves.