fix(json): do not treat booleans as numbers in JSONDiff - #224
Open
fei (feiiiiii5) wants to merge 1 commit into
Open
fei (feiiiiii5) wants to merge 1 commit into
fei (feiiiiii5) wants to merge 1 commit into
Conversation
`bool` is a subclass of `int` in Python, so the isinstance check that
selects the numeric scorer also matched booleans. `NumericDiff(True, 1)`
is then `1 - |1 - 1| / (1 + 1) = 1.0`, so a JSON boolean compared
against the integer 1 scored a perfect match:
JSONDiff().eval({"flag": True}, {"flag": 1}).score # 1.0
The TypeScript implementation tests `typeof v === "number"`, which is
false for booleans, so it never took that path and scored the same input
0.0. AGENTS.md states the two implementations share behavior; they did
not for booleans.
Excluding `bool` sends Python booleans to the same fallback the
TypeScript version reaches, and the two now agree on all seven cases I
checked against the published 0.3.0 packages, including the equal-boolean
cases (1.0) and the 'true' vs true string case (0.6667).
Test: `pytest py/autoevals/test_json.py -k booleans` fails on main with
`AssertionError: {'flag': True} scored equal to {'flag': 1}` and passes
here. The rest of the non-LLM suite passes; test_partial fails on main
too, for a missing OPENAI_API_KEY.
fei (feiiiiii5)
force-pushed
the
fix/json-diff-bool-is-not-number
branch
from
September 28, 2026 08:13
9a66d76 to
72e928d
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the case where a JSON boolean compared against an integer scores a perfect match.
JSONDiffpicks the numeric scorer with a bareisinstancecheck:boolis a subclass ofint, so booleans take that branch.NumericDiff(True, 1)is1 - |1 - 1| / (1 + 1) = 1.0, and the same happens for0vsFalse:For a scorer, a boolean and an integer are different JSON values, and a model that returned
truewhere1was expected is being marked fully correct.The TypeScript implementation tests
typeof v === "number", which is false for booleans, so it never entered that branch and scored the same input0.0.AGENTS.mdsays the two implementations share the same scorer behavior; here they did not.Excluding
boolfrom the numeric check sends Python booleans to the sameJSON.stringifyfallback the TypeScript version already reaches. Running the same inputs through both publishedautoevals0.3.0 packages, they now agree:{"flag": true}vs{"flag": 1}{"flag": false}vs{"flag": 0}[true]vs[1]{"flag": true}vs{"flag": true}{"flag": "true"}vs{"flag": true}The last two rows matter: equal booleans still score 1.0, and the string
"true"still does not match the boolean.Test:
pytest py/autoevals/test_json.py -k booleansfails onmainat9546b28withAssertionError: {'flag': True} scored equal to {'flag': 1}and passes on this branch.pytest py/autoevals/test_json.pyis 6 passed. The other non-LLM test files pass;test_partialfails onmainas well, from a missingOPENAI_API_KEYin this environment.black,ruffandcodespellare clean on both changed files at the versions in.pre-commit-config.yaml. I did not run the LLM-backed tests, which need credentials.The list branch divides by
max(len(o1), len(o2))while iteratingzip(o1, o2), so extra elements lower the score rather than being ignored. That is intended and is identical in both implementations, so I left it alone.