Data dictionary and annotation rules
September 18, 2026 · View on GitHub
The benchmark contains exactly two UTF-8 CSV files, joined by patient_id. Each has 1,000 data rows. Quoted progress-note cells contain embedded newlines. Numeric/date blanks mean null, not zero. Booleans are true/false. Read partial dates as strings.
Model inputs
| Field | Meaning |
|---|---|
patient_id | Unique synthetic identifier. |
age | Integer age at encounter, 50–80 inclusive. |
sex | Administrative female/male field, not gender identity. |
note_date | Encounter date, YYYY-MM-DD; anchor for relative time. |
progress_note | Full synthetic outpatient note. |
Answer key
patient_profiles.csv repeats the four demographic/identifier columns, followed by these ten scored fields:
| Field | Values / meaning |
|---|---|
smoking_status | current, former, never, unknown. |
pack_years | Point value for exact/approximate estimates; otherwise null. |
pack_years_min | Point value, lower range endpoint, or inclusive minimum; null if unknown. |
pack_years_max | Point value or upper range endpoint; null for unknown or lower-bound-only. |
pack_years_type | exact, approximate, range, lower_bound, unknown. Exact means unqualified in the note, not objectively measured. |
pack_years_basis | documented, calculated, never_zero, unknown. |
quit_date | YYYY, YYYY-MM, YYYY-MM-DD, or null. |
quit_date_precision | year, month, day, unknown, not_applicable. |
quit_date_basis | documented, derived_relative, unknown, not_applicable. |
quit_date_approximate | Boolean for an available date; otherwise null. |
Unscored audit columns are status_evidence, pack_years_evidence, quit_date_evidence, challenge_tags, and note_format. Evidence spans appear verbatim in the corresponding note; they may identify missing information or ambiguity. Tags describe construction scenarios; note_format identifies one of 12 formatting families. None of these audit columns is sent to either model.
Rules
Cigarette scope. Classify the patient's own cigarette history. Other people's smoking, passive exposure, vaping, cigars, cannabis, and chewing tobacco do not establish cigarette status or pack-years.
Status. Current includes daily/intermittent use and relapse. A future quit plan is current. Former requires prior use and present abstinence, including a recently completed quit. Never requires explicit lifetime denial. No history, unresolved contradiction, historical use without an update, and “nonsmoker” or “none currently” without enough lifetime context are unknown. Explicitly identified corrections override stale fields. Do not guess from medication or diagnosis.
This is clinical-documentation extraction, not a survey recode: an undocumented lifetime threshold of 100 cigarettes is not imposed. No less-than-100-cigarette experimentation cases are present. CDC's survey definitions are informative background, not identical to this benchmark policy.
Pack-years. Use 20 cigarettes per pack and sum packs/day × smoking years over supported periods. Exclude nonsmoking gaps. Both bounds equal the point for exact/approximate estimates. Preserve ranges without inventing a midpoint. An inclusive lower bound has a minimum but null point and maximum. Approximate points do not imply an invented uncertainty interval. Never-cigarette history yields zero. Incomplete exposure yields null. A stale cumulative total with explicitly unknown subsequent exposure does not establish today's lifetime total.
Quit dates. The target is the final cessation initiating current maintained abstinence. Ignore prior failed quits, future plans, others' histories, and cessation of other substances. Preserve date precision. Exactly N weeks before the encounter maps to encounter date minus 7N days. About N years ago maps to encounter year minus N with year precision and approximate=true; this is a declared normalization convention. Former with no date has unknown date fields. Current/never has null date, not_applicable precision/basis, and null approximate flag.
Answers represent information supported by the note; there are no hidden numeric/date truths that models must guess. The complete scoring policy supplied to both models is extraction_policy.txt.
Scoring
All ten fields must match for complete-extraction correctness. Numeric tolerance is ±0.1 pack-years for point and bounds; categories, nulls, date strings, provenance, precision, and approximate flags match exactly. Missing predictions count as incorrect. Smoking status also has macro-F1 and a four-class confusion matrix. Subgroups prevent never-smoker zeros and nonapplicable dates from dominating interpretation. Both normalized and unnormalized results are available.
Definition sources
These sources informed definitions; individual rows are synthetic, not sourced patient observations.