Pediatric medication orders are usually written by weight. A clinician records the child’s weight, a dose such as 15 mg/kg/day, how many times a day it is given, and the route, and the amount the child actually receives follows from those elements. That field arithmetic is simple. Reading the order correctly is not.
Anyone, or anything, that reads a clinical note and reconstructs the intended dose has several chances to go wrong: pick an old weight instead of the current one, confuse a daily dose with a per-dose figure, read a volume as a dose, forget a stated maximum, or multiply correctly from the wrong inputs. Dosing error is the largest category of pediatric medication error in the studies that have measured it, and tenfold errors traced to calculation and decimal mistakes have been documented over years at a single children’s hospital.
Large language models are now good enough at reading clinical text that it is tempting to hand them this job. Whether they reconstruct the regimen correctly, and whether the obvious engineering safeguard of taking the arithmetic away from them actually helps, are empirical questions. A pre-registered benchmark set out to answer both.
Key Takeaways
- Pediatric medication orders often rely on weight-based calculations, which can lead to errors during interpretation.
- Dosing errors rank as the largest category of pediatric medication errors, primarily due to incorrect weight usage and misinterpretation of dosage.
- A new benchmark tests language models on both extraction of medication information and their calculation accuracy simultaneously.
- Initial results show leading models perform well but still struggle with issues like outdated weights and ignoring stated dosage limits.
- This benchmark indicates the need for clear specifications and comprehensive checks before deploying language models in clinical settings.
Table of contents
- Why Extraction and Dosing Need to Be Tested Together
- How the Benchmark Was Built
- Two Workflows, One Question
- What the Leading Models Did
- Pattern One: Nobody Asked How Old a Weight May Be
- Pattern Two: The Maximum Was Read and Then Ignored
- Pattern Three: The Calculator Did Not Rescue the Weak Model
- What This Suggests for Deployment
- Status and Limits
- Looking Ahead
Why Extraction and Dosing Need to Be Tested Together

Most evaluations of language models on medication information test one thing at a time: whether a drug name or a dose was extracted, or whether the model can do a calculation when the inputs are given to it.
A weight-based regimen is different because its elements depend on each other. A correct weight and a correct mg/kg value still produce the wrong amount if the model treats a daily figure as a per-dose figure. A correct per-dose amount can still sit inside a wrongly described regimen, which matters to any downstream system that relies on the structured fields rather than the number.
The benchmark therefore scored four levels at once: whether each field was extracted; whether all required fields were correct together; whether the mg/kg values the model stated were consistent with the regimen in the note; and whether the resulting exposure departed from the intended regimen by more than 10 percent, or could not be reconstructed at all. The 10 percent figure adapts a threshold that a panel of pediatric critical care clinicians agreed on for what counts as a dosing error. The study applies it to a constructed intended dose and makes no claim about harm to a patient.
How the Benchmark Was Built
The test set is synthetic by construction. Each of its 361 records is a short clinical note for a child from birth to 18 years, with a weight drawn from the WHO and CDC growth references and a regimen for one of 21 medications whose pediatric dosing comes from the FDA label. Notes range from a bare three-line order to an admission note carrying a second medication, a discontinued order for the same drug, a weight from a prior visit, or a weight in pounds. Fiftyseven records are abstention cases, where a required element is missing and the correct answer is to say so rather than guess.
The design was registered on the Open Science Framework before any evaluated model saw a test record. The registration carries the protocol, the model configuration, and cryptographic hashes of the test set and the scoring code, so a reader can verify later that neither was changed after the fact. Every change made after registration is a dated entry in a public deviations log.
The benchmark was designed and run as a solo, self-funded project by Vinod Rufus Motani, an AI engineer and independent researcher based in the Chicago area, with no patient records, no institutional resources and no clinician panel, so that every number in it could be published with the data that produced it. Ground truth is known by construction, because each note was generated from a canonical regimen, and the author audited a blinded sample. The absence of a clinician or pharmacist review is stated as a limitation.
A simplified sequence is:
Synthetic note → Language model → Structured regimen → Scoring against the regimen the note was built from
Five models were specified in the registration and evaluated through their published interfaces: GPT-5.5, Claude Sonnet 5, Gemini 3.1 Pro, DeepSeek V4 Pro, and a small open-weight model, Llama 3.1 8B, run locally as the deliberate contrast. Four have complete results; validation of the fifth is in progress.
Two Workflows, One Question

Each model saw every record twice. The registered protocol also described a rules-only workflow with no language model; it was not run, the decision is recorded in the study’s public deviations log, and the two workflows below keep the letters B and C they carry in the registration and in every result file.
In the end-to-end workflow, the model extracted the fields, computed the per-dose amount itself and reported its own mg/kg values. In the hybrid workflow, the model extracted the fields only, and a deterministic calculator computed the exposure from them.
The hybrid design is widely assumed to be the safer one, on the reasoning that a calculator does not make arithmetic mistakes. The benchmark was built to test that assumption rather than to rank vendors.
What the Leading Models Did
On routine extraction, the three completed frontier models were at or near ceiling. Field accuracy was between 0.996 and 1.000, and complete dose-consistent regimen accuracy between 0.970 and 1.000, with no statistically significant pairwise difference among them after correction for multiple comparisons. Run-to-run agreement on a repeated subset was above 0.97 for every completed model.
The small model extracted individual fields almost as well, at 0.968 when computing the dose itself, but produced a fully consistent regimen in only one record in three.
Near-ceiling results for the strongest models were an expected outcome and were written into the protocol before the runs. The informative strata were pre-identified as the hardest notes, the abstention cases and the small model. That is where three patterns stand out.
Pattern One: Nobody Asked How Old a Weight May Be
All three completed frontier models, in both workflows, missed the same five of the 57 abstention cases in the same way.
Each of those five notes documents only a weight from a visit three months earlier. Each model used it as the current weight instead of flagging the gap. For an 8-month-old in that set, the prior weight is about 20 percent below the reference weight the generator withheld from the note.
The prompt gave no rule for how old a documented weight may be. The uniformity across vendors makes a gap in the specification the most likely explanation; whether an explicit rule changes the behaviour is a separate test the study has not yet run.
Pattern Two: The Maximum Was Read and Then Ignored
Five admission notes gave a weight-based rule, a maximum written into the order, and no per-dose amount, and in each the weight-based product exceeded the maximum.
Two models wrote the uncapped amount: 1478 mg of amoxicillin per dose against an 875 mg per-dose maximum stated in the order; 1874 mg of ceftriaxone per administration on a twice-daily order, 3748 mg a day, against a 2 g daily maximum; and 82 mg, 131 mg and 65 mg of prednisolone once daily against a 60 mg daily maximum. Six of the seven outputs overshoot by more than 10 percent and were scored as failures; the seventh, 65 mg against 60, overshoots by 8 percent and passed the tolerance, although the study’s separate maximum flag recorded it. A tolerance can pass what a limit check catches.
In every one of the seven outputs the model had placed the maximum in the correct structured field before computing the amount without it. The third frontier model applied the maximum in all five notes. This is not an extraction failure and not a specification gap; the instruction was in the note and the model read it.
An exploratory check applied after the fact, comparing each written amount with the maximum the model itself had extracted, flagged these seven end-to-end outputs and four hybrid-workflow outputs among the three frontier models, and nothing else from them. A retrospective check of this kind has not been tested as a prospective safeguard, but it shows that the information needed to catch the error was already in the model’s own output.
Pattern Three: The Calculator Did Not Rescue the Weak Model
The pre-registered question was whether handing the arithmetic to a calculator lowers the failure rate for the same model.
For one frontier model it did, with a confidence interval excluding zero; for another the reduction was small and its interval included zero; for the third there was nothing to reduce. Under the rule fixed in the protocol, which required the effect in a majority of models, the hypothesis is not supported on that measure. The protocol also states the comparison on a second measure, the primary endpoint, and on that wording two of the four completed models show a gain, so the verdict on that wording waits for the fifth model. Both are reported.
For the small model the calculator made the structured-regimen measure worse and changed the kind of failure. Failures rose from 196 to 271 of 304 records, and their composition moved from 189 inconsistent mg/kg values and 7 unreconstructable exposures to 2 and 269.
What that measure does not say is that the model’s written amount was wrong. Row by row, the small model copied the per-dose amount from the order correctly in about five records of six under both workflows and then described the regimen inconsistently. Computing the dose itself, it wrote the right amount in 154 records and then put the milligram figure or the per-day figure into the per-dose mg/kg field. With the calculator, it wrote the right amount in 233 records, labelled the order a fixed dose, and dropped the mg/kg value, leaving the calculator nothing to multiply.
A user reading only the amount would see a correct number most of the time. A system relying on the structured regimen, which is what a calculator or a dose-range check needs, would not.
What This Suggests for Deployment
A simplified sequence for anyone putting a language model in front of weight-based orders is:
Specify → Extract → Check completeness → Check limits → Compute → Review
Each step addresses a different failure. A missing exposure is the easier failure to catch, since a completeness check flags it. A plausible wrong number needs a range or maximum check, which would have caught the seven outputs above but would not catch a wrong weight. A stale weight needs a rule the prompt must state and a test that confirms the model follows it.
None of these checks is new. The benchmark’s contribution is to show, on a locked and public test set, which of them the strongest current models still need, and that the hybrid design does not substitute for them.
Status and Limits
Four of the five registered models have completed every run, including a 108-record subset repeated three times to measure stability. The fifth is under validation and its results are incorporated before publication.
Of the five pre-registered hypotheses, two are settled as not supported: that complete-regimen accuracy would trail field accuracy by at least ten points for every model, and that failures would rise with note complexity. The workflow hypothesis is not supported under one protocol wording and pending under the other. The hypothesis that confusion between daily and per-dose figures would dominate the errors is provisional until blinded manual coding is complete. The exploratory age-band hypothesis is inconclusive.
The records are synthetic, so the benchmark bounds performance on documentation it can fully specify rather than on the messier reality of an electronic health record. One test-set record has a printed amount more than 10 percent from the exact product of weight and dose, which no internally consistent answer can satisfy under both scoring conventions; it is disclosed and removed by the pre-registered sensitivity check. Two of the stated maxima in the headline notes are the author’s documented derivations from the labels rather than figures the labels state as such, and the article describes them as stated in the order for that reason.
Read Next
A few related pieces worth your time:
- The Evolving Role of the Municipal CIO in the Age of AI
- Humaniser’s Dual Advantage: Why the Best Free AI Humanizer Matters
- Ryne AI Content Evolution Method: From AI Drafts to Truly Human Writing
Looking Ahead
Language models read pediatric orders well. The patterns in this benchmark are not about reading. They are about what happens after reading: whether a model notices what the note does not say, whether it applies a limit it has already extracted, and whether it describes the regimen in a form that the next system can trust.
Those are the things to specify, test and check before deployment, and a locked, public benchmark is one way to find out which of them a given model still needs.
The benchmark’s code, synthetic data, locked test set, protocol, model outputs, analysis outputs, evidence packet, figures and supporting documents are available at https://github.com/VinodRufus/from-fields-to-dose.











