Evaluation · reliability
Reliable French legal AI: reliability must be demonstrated, not declared
Legal AI is reliable only for a defined use and level of risk. Test it on a closed question set, inspect every source and date, measure omissions, probe whether it asks for decisive facts and record discrepancies. No single score guarantees the next answer: reliability requires a reproducible protocol, ongoing monitoring and human review proportionate to the stakes.
Scope: French law · information, not legal advice
Reproducible reliability scorecard
Score observable dimensions separately: corpus, citation support, freshness, abstention, data governance and human control.
| Dimension | Repeatable test | Decision rule |
|---|---|---|
| Corpus coverage | Define the expected authorities and count answers that omit a relevant source. | Measure |
| Citation entailment | Open each reference and confirm that the cited passage supports the associated claim. | Measure |
| Legal freshness | Introduce known legislative changes and check whether the applicable version is used. | Measure |
| Abstention and clarification | Use incomplete questions; the system should request decisive facts or disclose uncertainty. | Measure |
| Data protection | Check stated purposes, retention, transfers, reuse and access controls. | Measure |
| Review and traceability | Record prompt, version, sources, expected result, discrepancy and reviewer decision. | Measure |
Choose a starting point without entering personal facts
Select a category only. No identity, date, amount, document or free-text question is included in analytics events.
Corpus coverage
Define the expected authorities and count answers that omit a relevant source.
Keep the output, verify it against the listed sources, then obtain human review for any material decision.
What is the most reliable AI for French law?
There is no absolute winner: reliability depends on the legal field, task, data, date and human control. Check the source, its date and the full document before using this point in a decision.
How do I measure legal hallucinations?
Count fabricated references, unsupported citations, outdated rules, invented facts and unsupported conclusions separately. Check the source, its date and the full document before using this point in a decision.
Is an answer reliable because it has sources?
No. A reference may exist while being irrelevant, outdated or unable to support the stated proposition. Check the source, its date and the full document before using this point in a decision.
How many questions should a benchmark contain?
Enough to cover real tasks and risk levels; disclose the set composition and scoring rules rather than relying on a headline number. Check the source, its date and the full document before using this point in a decision.
Should refusal behaviour be tested?
Yes. A system should request clarification, decline an unsupported conclusion and direct high-risk matters to human review. Check the source, its date and the full document before using this point in a decision.
How is reliability monitored after purchase?
Replay the same test set regularly and add real failures without recording unnecessary personal data. Check the source, its date and the full document before using this point in a decision.
Is confidentiality part of reliability?
Yes. An accurate output obtained through uncontrolled data processing remains unsuitable for the use case. Check the source, its date and the full document before using this point in a decision.
Can Julie guarantee zero errors?
No. Julie should be tested like any other tool: inspect sources, keep limits visible and use human validation according to risk. Check the source, its date and the full document before using this point in a decision.
I want to evaluate a French legal AI system. Help me design a reproducible benchmark for source support, freshness, omissions, abstention and confidentiality.