Skip to content

fix(json): do not treat booleans as numbers in JSONDiff - #224

Open
fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/json-diff-bool-is-not-number
Open

fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/json-diff-bool-is-not-number

Conversation

@feiiiiii5

Copy link
Copy Markdown

Fixes the case where a JSON boolean compared against an integer scores a perfect match.

JSONDiff picks the numeric scorer with a bare isinstance check:

elif (isinstance(o1, int) or isinstance(o1, float)) and (isinstance(o2, int) or isinstance(o2, float)):
    return self.number_scorer.eval(o1, o2).score

bool is a subclass of int, so booleans take that branch. NumericDiff(True, 1) is 1 - |1 - 1| / (1 + 1) = 1.0, and the same happens for 0 vs False:

JSONDiff().eval({"flag": True}, {"flag": 1}).score    # 1.0
JSONDiff().eval({"flag": False}, {"flag": 0}).score   # 1.0
JSONDiff().eval([True], [1]).score                     # 1.0

For a scorer, a boolean and an integer are different JSON values, and a model that returned true where 1 was expected is being marked fully correct.

The TypeScript implementation tests typeof v === "number", which is false for booleans, so it never entered that branch and scored the same input 0.0. AGENTS.md says the two implementations share the same scorer behavior; here they did not.

Excluding bool from the numeric check sends Python booleans to the same JSON.stringify fallback the TypeScript version already reaches. Running the same inputs through both published autoevals 0.3.0 packages, they now agree:

input before (py) before (ts) after
{"flag": true} vs {"flag": 1} 1.0 0.0 0.0
{"flag": false} vs {"flag": 0} 1.0 0.0 0.0
[true] vs [1] 1.0 0.0 0.0
{"flag": true} vs {"flag": true} 1.0 1.0 1.0
{"flag": "true"} vs {"flag": true} 0.6667 0.6667 0.6667

The last two rows matter: equal booleans still score 1.0, and the string "true" still does not match the boolean.

Test: pytest py/autoevals/test_json.py -k booleans fails on main at 9546b28 with AssertionError: {'flag': True} scored equal to {'flag': 1} and passes on this branch. pytest py/autoevals/test_json.py is 6 passed. The other non-LLM test files pass; test_partial fails on main as well, from a missing OPENAI_API_KEY in this environment. black, ruff and codespell are clean on both changed files at the versions in .pre-commit-config.yaml. I did not run the LLM-backed tests, which need credentials.

The list branch divides by max(len(o1), len(o2)) while iterating zip(o1, o2), so extra elements lower the score rather than being ignored. That is intended and is identical in both implementations, so I left it alone.

`bool` is a subclass of `int` in Python, so the isinstance check that
selects the numeric scorer also matched booleans. `NumericDiff(True, 1)`
is then `1 - |1 - 1| / (1 + 1) = 1.0`, so a JSON boolean compared
against the integer 1 scored a perfect match:

    JSONDiff().eval({"flag": True}, {"flag": 1}).score   # 1.0

The TypeScript implementation tests `typeof v === "number"`, which is
false for booleans, so it never took that path and scored the same input
0.0. AGENTS.md states the two implementations share behavior; they did
not for booleans.

Excluding `bool` sends Python booleans to the same fallback the
TypeScript version reaches, and the two now agree on all seven cases I
checked against the published 0.3.0 packages, including the equal-boolean
cases (1.0) and the 'true' vs true string case (0.6667).

Test: `pytest py/autoevals/test_json.py -k booleans` fails on main with
`AssertionError: {'flag': True} scored equal to {'flag': 1}` and passes
here. The rest of the non-LLM suite passes; test_partial fails on main
too, for a missing OPENAI_API_KEY.
@feiiiiii5
fei (feiiiiii5) force-pushed the fix/json-diff-bool-is-not-number branch from 9a66d76 to 72e928d Compare September 28, 2026 08:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant