Clinical AI safety tested from a practising physician perspective.
Independent clinical review, red-teaming and structured evaluation of healthcare AI outputs, with particular focus on high-risk perioperative, anaesthesia and acute-care scenarios.
What I evaluate
Practical, physician-led testing for teams building or deploying healthcare AI.
Clinical output review
Accuracy, safety, relevance, completeness and evidence alignment.
Red-team testing
Edge cases, escalation thresholds, contraindications and high-risk context.
Rubric development
Structured scoring criteria, severity taxonomies and critical-error definitions.
Patient-safety review
Unsafe reasoning, omissions, hallucinations, false reassurance and prioritisation errors.
Scenario design
Specialty-specific cases for perioperative medicine, anaesthesia and acute care.
Actionable remediation
Recommendations for prompts, guardrails, benchmarks and evaluation workflows.
A safety-first evaluation method
A single plausible serious-harm error can override an otherwise acceptable aggregate score.
Severity classification
Critical: plausible serious harm, delayed life-saving treatment or dangerous false reassurance.
Major: materially changes diagnosis or management or misses an important risk.
Minor: limited clinical consequence involving nuance, clarity or completeness.
Physician-led, clinically grounded
Practising anaesthesia physician in Australia with nearly 18 years of clinical experience across anaesthesia, perioperative medicine, cardiac anaesthesia, critical care and patient safety. Experience also includes medical-student teaching and clinical assessment.
Independent clinical evaluation and safety review only; not regulatory certification, legal advice or medical-device approval.