Workflow scorecards

Benchmarked by legal task, not generic leaderboard score.

QLLM reports task-specific accuracy, faithfulness, citation behavior, format adherence, latency, and cost. Open-ended legal outputs require lawyer review before external claims.

Latency caveat: the current live measurements are cold-start staging results through Modal-hosted backends. They are not warmed steady-state SLA numbers.

Initial live benchmark results

AreaNResultCold-start p95
Compliance concept detection21.000 accuracy155s
Clause diff classification20.500 accuracy180s
Contract type classification20.500 accuracy2.3s
RAG no-source refusal20.500 refusal behavior0.1s
Regular 70B summarization30.767 faithfulness precheck442s
Accuracy

Structured tasks

Classifiers are evaluated with exact labels, invalid output rates, and output contract adherence.

Faithfulness

RAG and summaries

Correctness-sensitive answers are scored against source context and queued for lawyer review.

Operations

Latency and cost

Cold-start and warm-start runs are tracked separately so commercial SLAs are not overstated.