Main Bank — Progress Log
The dominant reason is 981 questions where a row labelled with a distractor
Main Bank — Progress Log
Pipeline (each stage reproducible from the scripts in this folder)
| Stage | Script | Result |
|---|---|---|
| extract | audit_main.py |
2299 questions pulled out of the deployed app |
| dedup | dedup_scan.py → stage1_dedup.py |
276 duplicates removed, 19 duplicate IDs re-keyed → 2023 |
| mechanical | stage2_fix.py → stage2_final.py |
1679 stray correct-answer rows dropped, 1255 c[1] anchored, 4134 prefixes stripped → main_stage2.json |
| ship | stage2_final.py |
766 shippable, 1257 held back |
| author | realign2.py + hand_*.json → apply_hand.py |
+22 hand-written → 788 |
| build | stage3_build.py |
Main Bank — 803 verified.html |
Why 1250 are still held
The dominant reason is 981 questions where a row labelled with a distractor letter contains the correct answer’s reasoning. The row must be deleted, and because the remaining three rows then cover only three of four distractors, one genuinely new explanation has to be written per question.
- 981 x still holds the correct row
- 90 option count 4
- 76 x labels wrong
- 35 x filler text
- 32 x rows too short to trust
- 23 x rows = 5
- 11 option count 3
- 1 x rows = 20
- 1 x rows = 3
Per-question ID lists are in quarantine2.json.
Automation ceiling, measured
realign2.py scores every x row against every option using weights that favour
tokens unique to one option within that question. It recovers only 17% of the
held set (216 questions), because most rows do not name their option in
recoverable words. I tried plain IDF coverage and per-question discriminative
weighting; the ceiling is the same. Beyond ~20%, this needs reading.
Known data loss (upstream, not recoverable from this file)
sys is uniformly "Obstetrics" on all 2023 and topic is absent, so the
block picker can only use the six bank groupings. 556 of 773 IDs are opaque
(I-D004, E-D066) with no recoverable discipline. The earlier pipeline
discarded this; recovering it means re-deriving from the original source files.
Clinical errors found and corrected
-
APPLIED_BRIDGE_001— wrong key. DKA case, pH 7.21, HCO3 10, PaCO2 25. Winter’s formula gives an expected PaCO2 of 21-25, and the measured value of 25 sits inside it, so compensation is appropriate. The question was keyed D (concurrent respiratory alkalosis) while its ownwfield stated “compensation is appropriate - no concurrent disorder”. Re-keyed to B, withc[1]rebuilt to match. -
Winter’s formula stated wrongly elsewhere. A duplicate pair keyed the same compensation question two ways:
Expected PaCO2 = (HCO3 x 2) − 5 ± 2(correct) againstExpected PaCO2 = (1.5 x HCO3) + 8 ± 2(not a real formula). Both are in the held set awaiting rewrite.
Key errors are the reason every batch gets an independent arithmetic check before the explanations are written, not after.
Progress by block
| Block | Shippable | Held |
|---|---|---|
| Medicine & Allied | 440 | 485 |
| Surgery & Allied | 162 | 354 |
| Eye & ENT | 39 | 80 |
| Obstetrics & Gynaecology | 34 | 122 |
| Paediatrics | 39 | 134 |
| Basic Sciences | 74 | 74 |