fix(ragas): resolve the judge model through the library default - #226
Open
fei (feiiiiii5) wants to merge 1 commit into
Open
fei (feiiiiii5) wants to merge 1 commit into
fei (feiiiiii5) wants to merge 1 commit into
Conversation
_get_model is what every RAGAS scorer calls to resolve its judge model.
It fell back to DEFAULT_RAGAS_MODEL, a hardcoded "gpt-5-nano", when the
caller passed no model and had not called init(default_model=...):
init() # resets to the library defaults
get_default_model() # 'gpt-5-mini'
_get_model(None) # 'gpt-5-nano' <- not the library default
So a RAGAS scorer silently judged with a different model than every other
scorer in the library. The TypeScript implementation resolves through
getDefaultModel() and never had a separate fallback, so the same
evaluation produced different scores on the two implementations.
The constant was already labelled deprecated, by the commit that added
init(default_model=...) (braintrustdata#161): "This was previously 'gpt-5-mini' but now
defaults to the configured model." Only the first half of that change
landed -- the scorers stopped taking model=DEFAULT_RAGAS_MODEL as a
default argument, but _get_model still returned the constant.
_get_model now delegates to get_default_model(), which also removes the
manual _default_model_var lookup. The constant had no other reader, so
it and that import go with it.
Test: py/autoevals/test_ragas_default_model.py fails on main with
"assert 'gpt-5-nano' == 'gpt-5-mini'" and passes here. Excluding the
LLM-backed files, the suite goes from 25 to 28 passed; the 4 failures and
8 errors are unchanged and are all missing OPENAI_API_KEY or a missing
litellm.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The RAGAS scorers pick a different judge model from every other scorer in the library when nothing is configured.
_get_modelis what all seven RAGAS scorers call to resolve their model. It fell back to a hardcoded constant:So with no explicit
model=and noinit(default_model=...):Every RAGAS scorer is judged by a different model than every other scorer in the same library, with nothing printed. The comment above the constant already says the hardcoded fallback is deprecated. It was labelled that way by the commit that introduced
init(default_model=...)(#161), whose message says the change was "replacing the hardcoded gpt-4o default" — but only half of it landed. That commit moved the scorers offmodel=DEFAULT_RAGAS_MODELas a default argument while leaving_get_modelreturning the constant.The TypeScript implementation never had a separate fallback —
js/ragas.tsresolves throughgetDefaultModel()— so the same evaluation gets a different judge on the two implementations, and therefore different scores._get_modelnow delegates toget_default_model(). That also removes the manual_default_model_varlookup it was doing by hand.DEFAULT_RAGAS_MODELhad no other reader anywhere in the repo, so it goes with it, along with the now-unused import.Test:
py/autoevals/test_ragas_default_model.pyfails onmainat9546b28withAssertionError: assert 'gpt-5-nano' == 'gpt-5-mini'and passes here. It covers the unset case, an explicitmodel=, and a configuredinit(default_model=...), so the two paths that already worked stay pinned. Excluding the LLM-backed test files, the suite goes from 25 to 28 passed; the 4 failures and 8 errors are unchanged and all come from a missingOPENAI_API_KEYor a missinglitellmin this environment. I did not run the LLM-backed tests, which need credentials.black,ruffandcodespellare clean on both files at the versions in.pre-commit-config.yaml.This changes which model runs by default, so if you would rather keep
gpt-5-nanofor RAGAS the right place to decide that is adefault_modelfor the RAGAS group rather than a second fallback — but I did not want to make that call unilaterally, and the two implementations should not disagree either way.