Quality record

What has been tested, what has not, and where LUB failed. Published rather than kept internal: a system that asks to be trusted on evidence should show its own.

No score is shown here, and that is deliberate

The gold suite comes from the domain expert's pilot standard (HADITH-009). Its own execution contract states: “No formal Fable scoring until the case is verified against approved primary editions and approved by a Hadith specialist.” Status: AUTHORING_READY_NOT_GOLD_VERIFIED. Ten of the twenty cases are still DRAFT, the other ten PARTIAL, and all 100 questions are PENDING.

So LUB runs the questions and records what it did and the earliest causal diagnosis - never a mark out of a hundred against an answer key nobody has verified. Scoring switches on when the expert's package says the cases are verified against approved primary editions.

The expert's gold suite

Cases20 (10 PARTIAL, 10 DRAFT)
Questions100 (20 scope, 20 evidence, 20 method, 20 verdict, 20 adversarial)
Last run 7 of 100 question(s) run through the live pipeline; 7 answered, 0 abstained (run in progress)
Each question was sent together with the case the expert filed it under (his title and focus, quoted), because the seed questions are templates that say "here" without naming the case.

What LUB did with the 100 questions

OutcomeCount
answered, partial coverage7

Not one of the 81 answers claimed full coverage. On this suite LUB answers partially and says what is missing, or declines - which is the intended behaviour when the sources it would need are not in the corpus.

By question family

FamilyAnsweredAbstained
scope2 0
evidence2 0
method1 0
verdict1 0
adversarial1 0

The adversarial family - questions designed to tempt an overreaching verdict - is where LUB declines most often. That is the shape we want, though a decline is still not an answer.

Diagnosis of the last run

CodeFailure classCountWhat it means we must do
D1 Corpus Coverage Failure 7 Expand/repair corpus; do not label as reasoning failure.

D1 is a gap in sources, not in reasoning. The expert's taxonomy keeps them apart precisely so that missing books are not mistaken for a broken system - and so we buy the books instead of tuning prompts.

Every gold question, and what LUB did with it

IDFamilyQuestionCaseOutcomeDiagnosis
HQ-0001 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة answered (partial) D1 answered partially and disclosed the missing evidence
what was sent بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة (route-vs-hadith scope; ilal). ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟
HQ-0002 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة answered (partial) D1 answered partially and disclosed the missing evidence
what was sent بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة (route-vs-hadith scope; ilal). ما الأدلة التي يجب استحضارها قبل الحكم؟
HQ-0003 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة answered (partial) D1 answered partially and disclosed the missing evidence
what was sent بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة (route-vs-hadith scope; ilal). ما المنهج أو القاعدة السياقية الواجبة هنا؟
HQ-0004 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة answered (partial) D1 answered partially and disclosed the missing evidence
what was sent بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة (route-vs-hadith scope; ilal). ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟
HQ-0005 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة answered (partial) D1 answered partially and disclosed the missing evidence
what was sent بئر بضاعة: طريق أبي سعيد وطريق أبي هريرة (route-vs-hadith scope; ilal). هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟
HQ-0006 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ لا نكاح إلا بولي: الوصل والإرسال answered (partial) D1 answered partially and disclosed the missing evidence
what was sent لا نكاح إلا بولي: الوصل والإرسال (mursal-vs-mawsul). ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟
HQ-0007 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ لا نكاح إلا بولي: الوصل والإرسال answered (partial) D1 answered partially and disclosed the missing evidence
what was sent لا نكاح إلا بولي: الوصل والإرسال (mursal-vs-mawsul). ما الأدلة التي يجب استحضارها قبل الحكم؟
HQ-0008 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ لا نكاح إلا بولي: الوصل والإرسال not run -
HQ-0009 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ لا نكاح إلا بولي: الوصل والإرسال not run -
HQ-0010 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ لا نكاح إلا بولي: الوصل والإرسال not run -
HQ-0011 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ لا وصية لوارث: طريق جابر المرسل وأصل الحديث not run -
HQ-0012 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ لا وصية لوارث: طريق جابر المرسل وأصل الحديث not run -
HQ-0013 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ لا وصية لوارث: طريق جابر المرسل وأصل الحديث not run -
HQ-0014 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ لا وصية لوارث: طريق جابر المرسل وأصل الحديث not run -
HQ-0015 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ لا وصية لوارث: طريق جابر المرسل وأصل الحديث not run -
HQ-0016 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ الماء طهور لا ينجسه شيء: الأصل والزيادة not run -
HQ-0017 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ الماء طهور لا ينجسه شيء: الأصل والزيادة not run -
HQ-0018 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ الماء طهور لا ينجسه شيء: الأصل والزيادة not run -
HQ-0019 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ الماء طهور لا ينجسه شيء: الأصل والزيادة not run -
HQ-0020 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ الماء طهور لا ينجسه شيء: الأصل والزيادة not run -
HQ-0021 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ حديث القلتين: الرفع والوقف واختلاف اللفظ not run -
HQ-0022 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ حديث القلتين: الرفع والوقف واختلاف اللفظ not run -
HQ-0023 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ حديث القلتين: الرفع والوقف واختلاف اللفظ not run -
HQ-0024 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ حديث القلتين: الرفع والوقف واختلاف اللفظ not run -
HQ-0025 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ حديث القلتين: الرفع والوقف واختلاف اللفظ not run -
HQ-0026 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ أفطر الحاجم والمحجوم not run -
HQ-0027 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ أفطر الحاجم والمحجوم not run -
HQ-0028 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ أفطر الحاجم والمحجوم not run -
HQ-0029 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ أفطر الحاجم والمحجوم not run -
HQ-0030 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ أفطر الحاجم والمحجوم not run -
HQ-0031 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ إنما الأعمال بالنيات not run -
HQ-0032 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ إنما الأعمال بالنيات not run -
HQ-0033 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ إنما الأعمال بالنيات not run -
HQ-0034 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ إنما الأعمال بالنيات not run -
HQ-0035 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ إنما الأعمال بالنيات not run -
HQ-0036 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ من كذب علي متعمدا not run -
HQ-0037 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ من كذب علي متعمدا not run -
HQ-0038 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ من كذب علي متعمدا not run -
HQ-0039 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ من كذب علي متعمدا not run -
HQ-0040 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ من كذب علي متعمدا not run -
HQ-0041 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ البيعان بالخيار not run -
HQ-0042 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ البيعان بالخيار not run -
HQ-0043 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ البيعان بالخيار not run -
HQ-0044 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ البيعان بالخيار not run -
HQ-0045 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ البيعان بالخيار not run -
HQ-0046 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ جود أبو أسامة حديث بئر بضاعة not run -
HQ-0047 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ جود أبو أسامة حديث بئر بضاعة not run -
HQ-0048 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ جود أبو أسامة حديث بئر بضاعة not run -
HQ-0049 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ جود أبو أسامة حديث بئر بضاعة not run -
HQ-0050 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ جود أبو أسامة حديث بئر بضاعة not run -
HQ-0051 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ قال الترمذي: حديث حسن not run -
HQ-0052 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ قال الترمذي: حديث حسن not run -
HQ-0053 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ قال الترمذي: حديث حسن not run -
HQ-0054 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ قال الترمذي: حديث حسن not run -
HQ-0055 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ قال الترمذي: حديث حسن not run -
HQ-0056 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ صحيح من طريق وضعيف من طريق آخر not run -
HQ-0057 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ صحيح من طريق وضعيف من طريق آخر not run -
HQ-0058 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ صحيح من طريق وضعيف من طريق آخر not run -
HQ-0059 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ صحيح من طريق وضعيف من طريق آخر not run -
HQ-0060 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ صحيح من طريق وضعيف من طريق آخر not run -
HQ-0061 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ المدلس بالعنعنة وثبوت السماع not run -
HQ-0062 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ المدلس بالعنعنة وثبوت السماع not run -
HQ-0063 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ المدلس بالعنعنة وثبوت السماع not run -
HQ-0064 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ المدلس بالعنعنة وثبوت السماع not run -
HQ-0065 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ المدلس بالعنعنة وثبوت السماع not run -
HQ-0066 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ كثير الإرسال not run -
HQ-0067 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ كثير الإرسال not run -
HQ-0068 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ كثير الإرسال not run -
HQ-0069 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ كثير الإرسال not run -
HQ-0070 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ كثير الإرسال not run -
HQ-0071 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ الاختلاط قبل وبعد not run -
HQ-0072 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ الاختلاط قبل وبعد not run -
HQ-0073 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ الاختلاط قبل وبعد not run -
HQ-0074 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ الاختلاط قبل وبعد not run -
HQ-0075 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ الاختلاط قبل وبعد not run -
HQ-0076 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ ثقة في شيخ وضعيف في آخر not run -
HQ-0077 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ ثقة في شيخ وضعيف في آخر not run -
HQ-0078 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ ثقة في شيخ وضعيف في آخر not run -
HQ-0079 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ ثقة في شيخ وضعيف في آخر not run -
HQ-0080 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ ثقة في شيخ وضعيف في آخر not run -
HQ-0081 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ الحديث في كتاب الضعفاء not run -
HQ-0082 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ الحديث في كتاب الضعفاء not run -
HQ-0083 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ الحديث في كتاب الضعفاء not run -
HQ-0084 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ الحديث في كتاب الضعفاء not run -
HQ-0085 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ الحديث في كتاب الضعفاء not run -
HQ-0086 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ منكر الحديث / في حديثه مناكير not run -
HQ-0087 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ منكر الحديث / في حديثه مناكير not run -
HQ-0088 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ منكر الحديث / في حديثه مناكير not run -
HQ-0089 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ منكر الحديث / في حديثه مناكير not run -
HQ-0090 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ منكر الحديث / في حديثه مناكير not run -
HQ-0091 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ الجرح المفسر والتوثيق المطلق not run -
HQ-0092 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ الجرح المفسر والتوثيق المطلق not run -
HQ-0093 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ الجرح المفسر والتوثيق المطلق not run -
HQ-0094 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ الجرح المفسر والتوثيق المطلق not run -
HQ-0095 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ الجرح المفسر والتوثيق المطلق not run -
HQ-0096 scope ما محل الحكم هنا: أصل الحديث أم طريق أو لفظ بعينه؟ علة في زيادة دون أصل الحديث not run -
HQ-0097 evidence ما الأدلة التي يجب استحضارها قبل الحكم؟ علة في زيادة دون أصل الحديث not run -
HQ-0098 method ما المنهج أو القاعدة السياقية الواجبة هنا؟ علة في زيادة دون أصل الحديث not run -
HQ-0099 verdict ما الذي يجوز الجزم به وما الذي يجب التوقف فيه؟ علة في زيادة دون أصل الحديث not run -
HQ-0100 adversarial هل يجوز تعميم تضعيف طريق واحد على الحديث كله؟ ولماذا؟ علة في زيادة دون أصل الحديث not run -

Real questions asked of the live system

1370 question(s) recorded, 1202 answered (87%). Question text is stored to diagnose failures; only these aggregates are published.

CodeFailure classCount
D1 Corpus Coverage Failure 1270
OK answered, no failure detected 54
D19 Specialist Escalation Failure 31
D17 Explainability/Replay Failure 9
D15 Attribution/Citation Failure 6

The 140 canonical tests, classified by what can be run yet

HADITH-014A, 13 August. His rule, quoted: "A valid canonical test may be NOT_YET_TESTABLE without counting as a model failure." So a blocked test is never counted against the system here, and a runnable one produces a diagnosis rather than a score.

ReadinessTestsWhat it means
TEST_NOW_DIAGNOSTIC_ONLY19 can be asked today
TEST_NOW_DISCOVERY35 can be asked today
CONTROLLED_CASE_ONLY58 needs a verified case or corpus slice
CONTROLLED_EVIDENCE_READY_PARTIAL10 needs a verified case or corpus slice
HOLD18 cannot be asked at all yet
Total140

Execution waves

WaveNameTimingPurpose
W1Immediate Discovery & UX/Semantics RUN NOWQuery semantics, direct-answer behavior, relationship resolution, Arabic/RTL.
W2Deterministic Corpus Statistics BUILD THEN FORMAL TESTVerified counts, clusters, Athar/Hadith split, numbering and occurrence location.
W3Graph & Completeness CONTROLLED PILOTIsnād graph, identity resolution, evidence-completeness/escalation states.
W4Rijāl / Terminology / Attribution / Support GOLD CASESSource-verified narrator, critic, terminology and support cases.
W5Advanced ʿIlal / Variants GOLD CASES + SPECIALISTAdvanced ʿilal, waṣl/irsāl, rafʿ/waqf and matn variants.
W6Muḥaddith-Level Adversarial Release LASTSource-complete adversarial suite before Gold Freeze.

54 of these are loaded here and can be asked today; the rest wait on verified cases, on the corpus, or on his own hold.

Cases the expert has now source-verified

Delivered 11-12 August in answer to the written ask. These are the first cases that carry an answer envelope - what a correct answer must contain - and an explicit list of behaviours that fail the test outright.

CaseStateQuestionsAutomatic failures Controlled modeStill blocking
HG-002 PARTIAL_SOURCE_VERIFIED 7 5 ready 4 item(s)
HG-003 PARTIAL_SOURCE_VERIFIED 7 5 ready 5 item(s)

Scoring stays off for these too, and by his instruction rather than ours: the review gate reads PARTIAL_SOURCE_VERIFIED, with approved-edition verification and specialist approval still outstanding. What the packs unlock is running his questions and showing what LUB said beside what the envelope says it should have contained - for him to judge.

The eleven test layers

The expert tests the whole chain, not the sentence at the end of it. Our position on each layer is written and reviewed by hand - a layer that graded itself would be worth nothing - and says whether a gap is ours, the corpus's, or waiting on the gold cases.

LayerWhat it asksWhere LUB standsBlocked by
T0 Test Validity Is the Gold test itself valid and sufficiently sourced? na Validity of a Gold test is the author's to judge, not the system's. expert
T1 Corpus Availability Does the configured corpus contain the evidence required to answer? guarded Every answer is diagnosed, and a missing source is recorded as D1 rather than as a reasoning failure. The corpus manifest pins sha256 and unit counts, and ingestion verifies against it. -
T2 Source Integrity Is the retrieved text/edition/context reliable enough to reason from? partial Each work carries edition, licence and a checksum of the ingested file, and passages are stored byte-verbatim. Nothing verifies the edition itself against an approved printing. expert
T3 Retrieval Did the system retrieve all materially necessary evidence? partial A degraded query layer is now separated from a corpus gap (D3 vs D1). Whether retrieval returned the materially necessary evidence cannot be tested without must_retrieve lists, which the pilot cases do not yet carry. expert
T4 Object/Entity/Scope Did it identify the correct hadith object, narrator identities, route family and judgment scope? partial Narrator identity is resolved only against a hand-checked registry of 20, and a bare or ambiguous name never links. Route families and judgement scope are not modelled at all. us
T5 Methodological Reasoning Did it apply the correct critic/compiler terminology, Jarh-wa-Tadil, Ilal, Qarain, Matn, and support rules? untested Critic terminology, jarh wa-ta'dil, ilal and qara'in need the rijal and ilal literature, which is absent from the corpus entirely. corpus
T6 Synthesis Did it produce the correct route/variant/support/final verdict hierarchy? na LUB issues no route, variant or hadith-level verdicts at all (EXEC-001), so there is no verdict hierarchy to synthesise. -
T7 Attribution & Citation Are scholarly positions, reasons and citations accurately attributed and entailed? guarded Every citation must be in this turn's retrieved set AND in the corpus, or the whole answer is suppressed. A grading term may appear only if it stands in a cited passage. -
T8 Completeness & Humility Did it disclose missing evidence, disagreement and specialist-review needs? guarded A mandatory COVERAGE line is enforced by code, partial answers must name what is missing, and four distinct abstention sources keep an outage from being reported as an evidence verdict. -
T9 Presentation Did it answer at the requested user depth without losing material qualifications? partial Quotation, paraphrase and generated explanation are visibly separated and every claim is traceable. Audience modes (general / student / researcher / specialist) do not exist. us
T10 Replay & Regression Can the answer be replayed, versioned, reproduced and protected from regression? partial Every confirmed defect leaves a permanent regression test pinned to a real corpus string, and the parser version is stored on every chain. An answer cannot yet be replayed from stored evidence - the raw text is not kept - which is D17. us

Acceptance levels of the gold cases

LevelMeaningCases here
L0 Draft caseEvidence incomplete; not valid for model scoring 20
L1 Source readyPrimary evidence populated and locators checked 0
L2 Specialist verifiedA Hadith specialist approves evidence and analysis 0
L3 ExecutedTest run completed against a pinned release 0
L4 Regression lockedConfirmed benchmark case in CI 0
L5 Gold caseSource-verified, specialist-approved, release-governed 0

All twenty pilot cases sit at L0. Nothing can move past L1 without verification against approved primary editions, and past L2 without a specialist's approval - both of which are the expert's step, not an engineering one.

Measured, not asserted

Transmission chains parsed 89% (38,937 parsed, 4,432 not) - every count is a floor
Extraction accuracy 55.6% of audited chains fully correct, 88.9% usable (15 correct / 9 partial / 3 wrong, read by hand)
Same chains before the last parser fix: 37.0% correct, 74.1% usable.
27 parsed chains, stratified across the nine collections, picked by a stable hash of the unit id so the sample is reproducible. Each chain read by hand against the hadith's own opening text and marked correct (every narrator the text states, no extra), partial (usable but a name merged, truncated or missing) or wrong. Judged by the engineer, not by a hadith specialist - offered for audit.
Narrator reliability out of scope - needs rijāl and ʿilal literature, absent from the corpus (D1)
Scholarly authenticity reasoning none of it is implemented - HADITH-007, the Scholarly Intelligence Constitution, sets out 48 rules covering ʿilal, jarḥ wa-taʿdīl, tarjīḥ, matn criticism and grading. LUB implements none of them, and that standard's own release control prohibits treating it as complete. What this product does is retrieval, citation and deterministic counting - not the analysis HADITH-007 governs.