Before an AI suggestion touches a patient's analysis, you should be able to answer one question about it: where did that come from? If the tool can show you — this rubric, in this repertory, at this location, with this grade; this passage, in this book — AI-assisted repertorization is an accelerant for exactly the discipline classical training teaches. If it cannot, no fluency of output makes it safe to use. This article turns that principle into a practical audit: ten checks a practitioner can run on any AI homeopathy tool — including ours — in under an hour, before trusting it in the consulting room. It differs from our broader pieces on AI in remedy selection and the case-analysis workflow: those explain what the technology does; this one is about verifying it. As always, it is written for practitioners and students — the clinical decision, here more than anywhere, remains yours.
Why Audit at All? The Two Patterns of AI in Homeopathy
Every AI feature in homeopathic software resolves, on inspection, into one of two patterns.
The retrieval pattern. The AI is a faster index. It takes the patient's words — "a squeezing band around the head, better pressing hard" — and surfaces candidate rubrics from real repertories, with locations and grades attached; or it drafts a shortlist strictly from rubrics you confirmed. Every output is a pointer into sources that exist and can be opened. The practitioner's craft — choosing, weighing, verifying in the materia medica — is unchanged; only the page-turning got faster.
The oracle pattern. The AI is an answer machine. Case description in, remedy name out, reasoning opaque, references absent or decorative. However accurate it might sometimes be, it is unauditable — and an unauditable step inserted into a clinical method breaks the method's chain of accountability at its most important link.
The audit below exists to tell these two patterns apart in practice, feature by feature — because real products (again: including ours) are mixtures, and the question is never "is AI good or bad?" but "which of this tool's outputs can I trace, and does anything untraceable ever reach the analysis?"
The Ten-Point Audit
Run these in order; they escalate from mechanics to governance. A tool that fails 1–4 is unsafe regardless of the rest. Keep notes — they become your re-audit baseline.
1. Rubric traceability
Take three AI-suggested rubrics and click through. Does each open at its exact location in a named repertory — chapter, rubric, sub-rubric — with its remedy list and grades visible? A suggestion that cannot be opened is the red flag. In Similia, every rubric the AI proposes is a live link into the repertory it came from; that is the design bar to hold any tool to.
2. Source verification of quotations
Where the AI quotes or paraphrases materia medica, open the passage. Does the cited book contain it, in recognisable form, about that remedy? Check three. A single fabricated quotation fails the audit outright — invented scholarship is worse than none, because it manufactures false confidence exactly where verification should live.
3. The invented-rubric probe
Feed the tool a symptom you know the repertories handle awkwardly — an odd modern phrasing, a rare modality. The correct behaviours are: nearest genuine rubrics (labelled as nearest), or an honest "no close match." The failing behaviour is a perfectly-worded rubric that exists in no repertory. This single probe catches more oracle-pattern tools than any other check.
4. Grade fidelity
For two suggested rubrics, compare the grades the AI displays against the repertory's own entry — bold, italic, plain; and in Murphy's four-grade scheme, the fourth degree handled correctly. Grades carry the evidential weight of the whole method; a tool careless with grades will be careless with everything downstream. (New to why grades matter this much? The repertorisation guide covers the foundations.)
5. Confirmation gating
Trace one full path from suggestion to analysis. Can any rubric or remedy enter the repertorisation without an explicit practitioner action? The trustworthy answer is no: the AI proposes, you confirm, and only confirmed items are analysed. Auto-inserted suggestions — however good — mean the tool, not you, is quietly taking the case.
6. Uncertainty behaviour
Give it a thin case: two vague symptoms, nothing striking. Does the tool say the picture is insufficient, hedge visibly, or rank candidates with the same confident face it wore for a rich case? Systems that cannot express uncertainty transfer their overconfidence to their users — the precise opposite of what a junior prescriber needs.
7. Semantic search precision
Test the everyday-language mapping with five phrasings your patients actually use — including one in your consulting language if you practise outside English — and judge whether the returned rubrics genuinely match the sense, not just the keywords. Then check the reverse: does a precise classical phrase find its exact rubric first? Our semantic search guide explains what good looks like under the hood.
8. Privacy and data governance
Read the tool's terms with three questions: Is consultation/case text used to train models? Who processes the data, and under what legal basis (GDPR processor terms, in Europe)? Can you export and delete? For recorded consultations, confirm the consent obligation sits where it belongs — with you — and that the vendor's handling supports it. A clinically brilliant tool with evasive data terms fails the audit; patient narrative is the most sensitive data a homeopath holds.
9. Model-change transparency
Ask (or find in the changelog): does the vendor announce when the AI's models or sources materially change? Continuous deployment means the tool you audited in January is not the tool you are using in July. You cannot re-audit what you do not know has changed — hence the twice-yearly re-audit rhythm in the FAQ.
10. The sign-off rule
Finally, audit yourself. Institute one personal rule: no AI-touched analysis is finished until the leading candidates have been read against the patient in the materia medica. Retrieval can be delegated; judgement cannot. Write the rule into your case template if that helps — the discipline costs three minutes and is the difference between AI-assisted practice and AI-directed practice.
Scoring It
There is no points system worth inventing: checks 1–5 are pass/fail (traceability, sources, no inventions, grade fidelity, gating), and a tool must pass all five before its conveniences deserve consideration. Checks 6–10 grade the maturity of the tool and of your governance around it. A tool passing all ten is not "AI you can trust blindly" — it is AI you never have to trust blindly, which is the actual goal.
What This Looks Like When It Works
Run honestly, the audit does not just filter tools — it changes how you use the one you keep. Suggestions become leads, not verdicts. The click-through to the rubric becomes reflex. The materia medica reading at the end stops feeling like a formality because it visibly catches the AI's near-misses — the rubric that matched the words but not the sense, the remedy strong in the grid but wrong in the gestalt. Practitioners who work this way report the same thing our own usage data shows: the AI's real gift is not answers but coverage — the third and fourth candidate rubric you would not have had time to hunt down by hand — while the decision quality still comes from where it always came from. That is the standard we build Similia's AI features to, and the audit above is how you hold us — and anyone else — to it.





