Form 1040-AI

Column Tax—Benchmark of Frontier Models on Tax Calculation

TaxCalcBench Leaderboard

arXiv No. 2507.16126
Model Use Only—Do not write or hallucinate in this space.

For the tax year Jan. 1–Dec. 31, 2025, or other benchmark edition beginning below

See separate instructions.
Tax year edition
Benchmark identification number TCB | 25 | v2
Publisher Column Tax
Home address (repository). If you have a fork, see instructions. github.com/column-tax/tax-calc-bench
Branch main
Returns tested 50
State CA IL NY VA
Models tested 38

Model Submission Campaign

Check here if you, or your model, want to be benchmarked. Email team@columntax.com. Checking a box below will not change your score.

Ranking Status

Check only one box.

Rank models by

Web Search

At any time during the benchmark, did the model (a) use a web search tool, or (b) otherwise look up tax forms and instructions online? (See instructions.) Show filers who answered

Show models

Providers

(see instructions)

If more than four providers, see below and check here

Results

Attach model outputs here. Also attach Schedule M for any line.

If your model is not listed, see instructions.

Click any line to open Sch. M.

TaxCalcBench leaderboard
For Disclosure, Privacy Act, and Paperwork Reduction Act Notice, see separate instructions. Cat. No. 2507161B Form 1040-AI (2025)
Form 1040-AI (2025) Page 2

Refund

Amount You Owe

Third Party Designee

Do you want to run this benchmark on your own model? See instructions.

Yes. Complete below.  No
Phone no.uv sync --all-extras
License (PIN)MIT

Sign Here

Joint return? See instructions. Keep a copy for your records.

Under penalties of perjury, I declare that I have examined this leaderboard and accompanying schedules and statements, and to the best of my knowledge and belief, they are true, correct, and complete. Declaration of model (other than taxpayer) is based on all tokens of which the model has any knowledge.

Your signatureColumn Tax
Date10/01/2026
Your occupationTax engine builders
If the model refused to prepare a return, enter the reason here (see inst.)
Model's signature. If a joint return, both must sign.The Models
Date10/01/2026
Model's occupationTax preparer (in training)
Email addressteam@columntax.com

Paid Preparer Use Only

Preparer's nameColumn Tax
Preparer's signatureColumn Tax
Firm's nameColumn Tax

2025

Instructions for Form 1040-AI

TaxCalcBench Leaderboard

What's New

Tax Year 2025 (v2 edition). The 2025 edition is harder than the 2024 edition in three ways: inputs are realistic PDF documents (W-2s, 1099s, and so on) instead of structured data, cases include state returns as well as federal returns, and the cases cover much more complex tax and financial situations.

Tax Year 2024 (v1 edition). Federal-only returns for relatively simple situations, with inputs provided as structured JSON. Each case was run four times and the scores were averaged (pass@1). Use the tax year boxes at the top of the form to switch editions.

General Instructions

What is TaxCalcBench?

TaxCalcBench measures whether frontier AI models can do the calculation step of tax preparation: given everything about a taxpayer, produce the completed return. Each test case pairs a taxpayer's inputs with the expected, correctly computed return, and the model's answer is compared to it line by line.

Tax calculation has traditionally been done by hand-built, deterministic tax engines that encode tens of thousands of pages of rules. This benchmark asks whether a model can do the same job on its own.

Who must file

Model providers who want their model on this form can email team@columntax.com. Anyone can also run the open-source harness from the repository.

Line Instructions

Column (a)—Correct returns (strict). The share of test cases where every evaluated line exactly matches the expected return. This is the number that matters: a tax return has to be completely correct to be filed.

Column (b)—Correct returns (lenient). The share of cases where every evaluated line is within ±$5 of the expected value. Many small misses come from computing tax with bracket math instead of the official tax tables.

Column (c)—Correct (by line). The average percentage of evaluated lines that exactly match. A single early mistake can cascade through the rest of a return, so this is usually much higher than column (a).

Column (d)—Correct (by line, lenient). The same as column (c), counting lines within ±$5 as correct.

Column (e)—Cost per return. Average API cost to produce one return. A blank means cost wasn't available; blanks are never treated as $0. 2025 edition only.

Column (f)—Time per return. Average generation time for one return. 2025 edition only.

Thinking level. Each model is tested across its supported reasoning levels, and each column reports the best setting for that column. Two numbers on the same line can come from different settings.

Web search. Lines marked Web search had a web search tool available, so the model could look up current forms and instructions.

Partial coverage. Lines marked with a case count, such as 40/50, are scored only on the cases that completed.

Specific Instructions

Model-specific notes are attached to each model's Schedule M. Click a line on page 1 to open it.

Paperwork Reduction Act Notice

We ask for the information on this form to find out whether a model can do your taxes. Models are not required to provide the information requested unless they want to be on the leaderboard. The average time burden for completing this form varies by model; see column (f). If you have suggestions for making this form simpler, we would be happy to hear from you at team@columntax.com.

Results synced from the TaxCalcBench repository on 10/01/2026.