NREHub

Main Bank — Progress Log

The dominant reason is 981 questions where a row labelled with a distractor

549 words ~2 min

Main Bank — Progress Log

Pipeline (each stage reproducible from the scripts in this folder)

Stage Script Result
extract audit_main.py 2299 questions pulled out of the deployed app
dedup dedup_scan.py → stage1_dedup.py 276 duplicates removed, 19 duplicate IDs re-keyed → 2023
mechanical stage2_fix.py → stage2_final.py 1679 stray correct-answer rows dropped, 1255 c[1] anchored, 4134 prefixes stripped → main_stage2.json
ship stage2_final.py 766 shippable, 1257 held back
author realign2.py + hand_*.json → apply_hand.py +22 hand-written → 788
build stage3_build.py Main Bank — 803 verified.html

Why 1250 are still held

The dominant reason is 981 questions where a row labelled with a distractor letter contains the correct answer’s reasoning. The row must be deleted, and because the remaining three rows then cover only three of four distractors, one genuinely new explanation has to be written per question.

  • 981 x still holds the correct row
  • 90 option count 4
  • 76 x labels wrong
  • 35 x filler text
  • 32 x rows too short to trust
  • 23 x rows = 5
  • 11 option count 3
  • 1 x rows = 20
  • 1 x rows = 3

Per-question ID lists are in quarantine2.json.

Automation ceiling, measured

realign2.py scores every x row against every option using weights that favour tokens unique to one option within that question. It recovers only 17% of the held set (216 questions), because most rows do not name their option in recoverable words. I tried plain IDF coverage and per-question discriminative weighting; the ceiling is the same. Beyond ~20%, this needs reading.

Known data loss (upstream, not recoverable from this file)

sys is uniformly "Obstetrics" on all 2023 and topic is absent, so the block picker can only use the six bank groupings. 556 of 773 IDs are opaque (I-D004, E-D066) with no recoverable discipline. The earlier pipeline discarded this; recovering it means re-deriving from the original source files.

Clinical errors found and corrected

  1. APPLIED_BRIDGE_001 — wrong key. DKA case, pH 7.21, HCO3 10, PaCO2 25. Winter’s formula gives an expected PaCO2 of 21-25, and the measured value of 25 sits inside it, so compensation is appropriate. The question was keyed D (concurrent respiratory alkalosis) while its own w field stated “compensation is appropriate - no concurrent disorder”. Re-keyed to B, with c[1] rebuilt to match.

  2. Winter’s formula stated wrongly elsewhere. A duplicate pair keyed the same compensation question two ways: Expected PaCO2 = (HCO3 x 2) − 5 ± 2 (correct) against Expected PaCO2 = (1.5 x HCO3) + 8 ± 2 (not a real formula). Both are in the held set awaiting rewrite.

Key errors are the reason every batch gets an independent arithmetic check before the explanations are written, not after.

Progress by block

Block Shippable Held
Medicine & Allied 440 485
Surgery & Allied 162 354
Eye & ENT 39 80
Obstetrics & Gynaecology 34 122
Paediatrics 39 134
Basic Sciences 74 74