LongTableBench: Benchmarking Long-Context Table Reasoning Across Real-World Formats and Domains

November 4, 2025 ยท View on GitHub

Official repository for LongTableBench (EMNLP'25 paper), a comprehensive benchmark for evaluating long-context reasoning over semi-structured tables across diverse formats, tasks, and domains.


๐ŸŒŸ Overview

LongTableBench is the first multitask benchmark designed to evaluate the reasoning ability of large language models (LLMs) over long-context semi-structured tables. It features diverse tasks, formats, and domains, ensuring comprehensive coverage of real-world reasoning challenges.

Key Features

  • 5,950 QA instances derived from 850 seed questions
  • 7 table formats: Markdown, HTML, JSON, LaTeX, SQL, XML, CSV
  • 18 domains, covering medical, finance, education, entertainment, etc.
  • Context lengths up to 128K tokens
  • Single-turn & multi-turn, single-table & multi-table scenarios
  • Rigorous symbolic verification, cross-model validation, and human review

๐Ÿงฉ Task Overview

LongTableBench includes six carefully designed tasks, evaluating three fundamental dimensions: Structural Complexity, Long-Range Dependencies, and Semantic Integration.

TaskAbbr.DescriptionPrimary Challenge
Exact MatchingEMLocate and extract exact cell values while maintaining structural awareness.Structural
Basic Conditional FilteringBCFApply simple logical filters (e.g., select rows by condition).Structural
Fuzzy Conditional ManipulationFCMHandle sorting, aggregation, and approximate matching under fuzzy conditions.Long-Range
Fact RetrievalFRRetrieve or verify facts grounded in tabular evidence.Long-Range
External Knowledge FusionEKFCombine tabular data with external or commonsense knowledge.Semantic
Irregular Numerical InterpretationINIInterpret unconventional numerical formats (Roman numerals, date variants, etc.).Semantic

Each task can appear in single-turn / multi-turn and single-table / multi-table settings.


๐Ÿ“ˆ Dataset Statistics

  • Total instances: 5,950
  • Length distribution:
    • 40% short (0โ€“8K tokens)
    • 35% medium (8Kโ€“32K)
    • 25% long (32Kโ€“128K)
  • Task proportion:
    • 35% Structural (EM, BCF)
    • 35% Long-range (FCM, FR)
    • 30% Semantic (EKF, INI)


๐Ÿงพ Benchmark Results

๐Ÿ“‹ Complete Results Table (Click to expand)
ModelsEMBCFFCMFREKFINIAvg.MDHTMLJSONSQLXMLLaTeXCSVAvg.Range8K32K128K
Llama-3.2-3B-Instruct24.3014.7813.762.9911.999.6813.9414.3513.3611.7216.1013.1412.3515.2313.7531.8321.7611.928.15
Qwen2.5-3B-Instruct17.147.079.4033.7436.4418.6320.4021.0420.1922.4624.6621.0020.1119.2521.2525.4927.2018.2115.78
Phi-3-mini-128k-instruct28.669.748.2933.3419.0517.3920.2520.1221.7522.6520.8320.7819.5216.5420.3130.0925.2621.2414.24
Phi-4-mini-instruct28.366.947.4122.8521.2213.6717.4918.2217.1519.9417.6119.6015.3313.6317.3536.3521.5519.0911.82
Qwen3-4B23.1511.079.8824.5914.8013.3016.7117.6916.7716.9417.5117.0017.3214.6416.8418.1021.7015.6612.77
Gemma-3-4B-it31.6013.4014.2921.9819.0819.9720.5621.9823.2418.9220.8720.3022.2920.7421.1920.3728.6719.9513.07
Llama-3.1-8B-Instruct47.7319.0921.2915.1122.4122.9725.6926.9424.6822.0727.5125.2926.5628.8325.9826.0434.6622.6619.74
Qwen2.5-7B-Instruct35.8515.1820.5850.7839.1130.5032.0531.9335.3036.8932.4334.7433.7230.5533.6518.8240.8429.3525.95
Qwen2.5-7B-Instruct-1M47.5220.1323.7952.7724.8824.6533.6134.8835.2935.7734.4435.6433.7331.5334.4712.3140.5532.4327.83
TableGPT2-7B41.1017.6818.8045.3439.8228.6432.2634.4632.5234.1032.4036.7333.3330.4833.4318.7042.1529.1125.50
TableLLM-Qwen2-7B18.004.595.3410.516.695.948.9810.558.638.1810.957.529.528.879.1737.3913.468.764.73
TableLLM-Llama3.1-8B27.721.872.3811.655.157.7810.3111.1210.8110.5310.7610.3910.507.2410.1938.0611.8812.146.90
TableLlama3.361.381.412.251.610.721.833.250.470.892.510.352.801.431.67173.234.970.360.17
Mistral-7B-Instruct-v0.329.3214.1615.5919.8915.4614.8519.1920.5619.6716.8118.9719.0618.5620.6919.1920.2729.1519.478.95
Ministral-8B-Instruct-241023.989.3714.8225.6112.7014.2517.2618.7217.4814.5318.2615.3817.1319.6817.3129.7332.9017.331.55
Qwen3-8B35.9714.1817.0246.6827.8725.2928.5529.5428.1128.7630.6527.4629.3229.0628.9911.0139.7325.1120.80
GLM-4-9B-Chat46.2518.8218.7932.8621.7420.7227.9227.8429.9527.1228.5231.2029.1327.2228.7114.2134.7927.0721.89
GLM-4-9B-Chat-1M46.8117.7917.6931.2121.7820.8727.2428.6427.7627.0029.4927.2630.1127.0528.1911.0332.7026.9522.09
GLM-4-9B-041431.9017.7420.1735.6128.1921.7826.5128.4127.4623.8927.1124.8228.6430.0927.2022.7744.6123.5811.35
Gemma-3-12B-it46.3421.5127.4846.7742.9029.9336.9039.4037.4734.0239.8434.6739.6137.9537.5715.4949.8334.7126.16
Mistral-Nemo-Instruct-240735.3218.0218.3049.1053.6932.4034.9936.2933.4733.7335.3634.8237.2037.1935.4410.5252.1335.0017.85
Qwen2.5-14B-Instruct47.2620.1726.6254.3442.8634.9338.2540.9841.6437.0939.2339.9341.7739.3239.9911.7151.7534.5828.43
Qwen2.5-14B-Instruct-1M56.5122.1629.3963.8437.3737.1542.3743.9344.9439.1444.6544.2746.2043.8043.8516.1250.9840.1136.03
Qwen3-14B53.1822.1325.8958.4952.6940.1642.6246.6645.3638.7846.7445.1145.7642.5044.4217.9354.5140.5432.81
Mistral-Small-3.1-24B-Instruct-250362.3328.9829.0141.9348.9542.4043.0646.6644.9841.3748.0244.5948.3045.4745.6315.1956.6041.2031.38
Qwen3-30B-A3B48.7316.5422.6755.5344.9133.6637.7440.1038.8536.4038.4537.8641.9539.7839.0514.2151.4035.3126.50
Qwen3-32B51.4824.6226.9251.3746.1942.6740.7244.2742.7840.7943.2845.5743.6443.8243.4510.9952.0937.7932.29
GLM-4-32B-041443.7125.4326.3955.3951.1435.2440.4847.2139.6235.6141.8836.3944.5145.7941.5727.9158.0943.0420.31
Llama-3.1-70B-Instruct62.5831.9029.1556.6936.2637.6744.3549.8546.1439.7446.1645.1846.1645.7847.997.7352.6246.8433.57
Llama-3.3-70B-Instruct60.3333.3733.4741.3141.6737.1342.7447.8841.4935.8346.8738.5749.1148.7244.0730.1256.6540.9630.60
Qwen2.5-72B-Instruct59.4332.1533.7263.5557.1745.9649.9452.6352.1645.1051.7452.5752.2353.5351.4216.4063.9746.5139.36
Mistral-Large-Instruct-241155.9629.6929.1057.1046.9944.2044.9949.8445.3740.1949.1342.2150.5549.3246.6622.2063.8745.2325.89
DeepSeek-V369.6344.5143.6566.3657.4454.1857.0961.3962.3753.6762.4258.9661.7159.7160.0314.5770.8054.6745.80
GPT-4o-mini-2024-07-1860.4331.7333.1749.5535.2636.1042.7641.7840.8537.5343.1941.0441.4641.4241.0413.7954.4438.6135.22
Gemini-2.0-flash71.1947.5142.1650.0560.7653.3955.0358.6657.6958.1458.9558.2057.3358.4458.202.7965.1355.2544.72
GPT-4o-2024-08-0678.9552.4049.9772.7668.7962.1165.6067.3664.4458.6765.6260.9965.8966.1864.1613.5378.5763.3354.92
Claude-3.5-sonnet-2024102276.4746.1342.2668.4565.3758.9160.6364.0960.7052.0358.8859.0561.1661.2659.6020.2475.5958.4647.84
Phi-4-mini-reasoning3.452.391.3511.793.104.514.255.183.953.744.735.203.963.364.3042.789.122.471.15
Qwen3-4B-thinking30.857.068.6221.9624.7819.8718.9218.2617.7519.3721.3522.8920.3116.9519.5530.3828.5615.4912.70
DeepSeek-R1-Distill-Qwen-7B7.964.004.5213.7813.136.618.288.995.505.587.515.638.227.737.0249.7217.265.062.52
DeepSeek-R1-Distill-Llama-8B33.2525.2415.1938.0733.5820.5329.2830.7130.1223.3828.2827.3434.7528.5429.0239.2039.9531.8716.02
Qwen3-8B-thinking34.208.6110.8632.1829.9724.0423.3224.9424.8222.8724.5127.0324.2322.2924.3819.4334.0018.8317.13
GLM-Z1-9B-041442.2733.9720.5344.8846.1930.3937.8942.5231.7938.4141.6333.8641.4940.3838.5827.8159.5739.2514.84
Qwen3-14B-thinking46.6614.5418.3851.1042.9931.7235.3138.0535.6432.3937.8838.3736.8035.6436.3916.4347.9430.4027.60
Qwen3-30B-A3B-thinking40.1011.9116.2837.8635.3627.1628.6329.8129.4029.2530.8631.4728.9427.9429.6711.8842.7924.0519.04
QwQ-32B59.4440.0728.7359.3358.6845.9350.3154.1254.0447.3953.5748.5853.9952.8052.0712.9165.4349.7235.77
DeepSeek-R1-Distill-Qwen-32B50.7339.0625.4352.4857.1644.5545.6052.1346.5241.3349.5844.1249.7150.4547.6922.6666.1347.7622.91
Qwen3-32B-thinking48.2914.8617.7048.9041.8034.0834.9438.9536.0930.9436.5638.5835.9436.8436.2722.0848.0930.1526.57
GLM-Z1-32B-041447.7037.5025.1339.4557.0234.6741.8143.6138.2644.9042.2339.8445.6942.7142.4617.4958.7441.1125.59
DeepSeek-R1-Distill-Llama-70B56.0242.1027.5751.8752.6144.5847.1751.2046.9043.4253.2043.8252.0851.0748.8120.0362.8849.6428.98
DeepSeek-R171.2466.2142.4162.6865.2857.8762.8267.9267.8461.2666.5564.2165.8165.2965.5510.1574.6364.3749.45
Gemini-2.0-flash-thinking-exp-01-2169.9554.7439.7651.4864.2654.0456.9858.6161.4461.2560.6959.5959.4760.1360.174.7070.1656.0544.74

Complete results available in our paper.


๐Ÿ“ Dataset Format

The dataset is organized as follows:

datasets/
โ”œโ”€โ”€ tables/          # Table files (by source and length)
โ””โ”€โ”€ questions/       # QA files (organized by length & turn type)

Single-Turn Format

{
  "question_id": "Unique ID",
  "db_id": "Database ID",
  "question": "Question text (possibly with external knowledge)",
  "answer": "Ground truth (list/dict format)",
  "highlighted_table": ["Relevant table IDs"],
  "is_multi_table": true,
  "question_type": "EM | BCF | FCM | FR | EKF | INI",
  "db_path": "Path to the source table"
}

Multi-Turn Format

{
  "question_id": "Unique ID",
  "db_id": "Database ID",
  "evidence": "Optional external knowledge",
  "QA": [
    {"round": 1, "question": "Q1", "answer": "A1"},
    {"round": 2, "question": "Q2", "answer": "A2"}
  ],
  "highlighted_table": ["Relevant tables"],
  "is_multi_table": true,
  "question_type": "EM | BCF | FCM | FR | EKF | INI",
  "db_path": "Path to the table"
}

๐Ÿš€ Quick Start

Prerequisites

git clone https://github.com/liyaooi/LongTableBench
cd LongTableBench
pip install -r requirements.txt

1. Model Deployment (vLLM)

vllm serve Qwen/Qwen2.5-7B-Instruct \
  --api-key $YOUR_API_KEY \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.95 \
  --max_model_len 131072 \
  --trust-remote-code

Note: You can modify pred.py to use other serving frameworks

2. Running Evaluation

python pred.py \
  --json-path datasets/questions/8k/longtablebench_single.json \
  --tables_path datasets/tables/ \
  --content_type single \
  --table_format markdown \
  --model-path Qwen/Qwen2.5-7B-Instruct \
  --output-path results/ \
  --key $YOUR_API_KEY

Note: You can also refer to pred.sh

Command Line Options

ParameterDescriptionDefault
--json-pathInput JSON fileRequired
--tables_pathDirectory containing tablesRequired
--content_typesingle or multi turnRequired
--table_formatTable formatmarkdown
--model-pathModel ID or pathRequired
--output-pathOutput directoryRequired
--keyAPI keyRequired
--urlCustom model API endpointNone
--num-gpus-totalTotal GPUs1
--num-gpus-per-modelGPUs per job1

Model Calling Mechanism:

  • Default implementation uses the openai library for model inference
  • Can be adapted to other frameworks by modifying pred.py
  • Current implementation supports both single-turn and multi-turn conversations

๐Ÿงช Evaluation Protocol

  • Metric: F1 score (macro over structured answers)
  • Setting: Zero-shot, greedy decoding (temperature=0)
  • Truncation: Middle truncation for inputs > context window
  • FR task: Must include evidence (table cell or row reference)

Note: The files in the category directory are required for proper functionality. If you modify their paths, ensure corresponding updates in eval/result_process.py to avoid errors.


๐Ÿ“œ License


๐Ÿ“š Citation

If you use this benchmark in your research, please cite:

@inproceedings{li-etal-2025-longtablebench,
    title = "{L}ong{T}able{B}ench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domains",
    author = "Li, Liyao  and  Tian, Jiaming  and  Chen, Hao  and  Ye, Wentao  and  Ye, Chao  and  Wang, Haobo  and  Wang, Ningtao  and  Fu, Xing  and  Chen, Gang  and  Zhao, Junbo",
    editor = "Christodoulopoulos, Christos  and  Chakraborty, Tanmoy  and  Rose, Carolyn  and  Peng, Violet",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",
    month = nov,
    year = "2025",
    address = "Suzhou, China",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-emnlp.638/",
    pages = "11927--11965",
    ISBN = "979-8-89176-335-7"
}

๐Ÿ“ Notes

  • Our code is constantly being optimized, which may be slightly different from the one in the paper.
  • For questions or issues, please open an issue on GitHub.