Structured tasks
Classifiers are evaluated with exact labels, invalid output rates, and output contract adherence.
QLLM reports task-specific accuracy, faithfulness, citation behavior, format adherence, latency, and cost. Open-ended legal outputs require lawyer review before external claims.
| Area | N | Result | Cold-start p95 |
|---|---|---|---|
| Compliance concept detection | 2 | 1.000 accuracy | 155s |
| Clause diff classification | 2 | 0.500 accuracy | 180s |
| Contract type classification | 2 | 0.500 accuracy | 2.3s |
| RAG no-source refusal | 2 | 0.500 refusal behavior | 0.1s |
| Regular 70B summarization | 3 | 0.767 faithfulness precheck | 442s |
Classifiers are evaluated with exact labels, invalid output rates, and output contract adherence.
Correctness-sensitive answers are scored against source context and queued for lawyer review.
Cold-start and warm-start runs are tracked separately so commercial SLAs are not overstated.