Skip to content

fix(ragas): score 0 when Faithfulness gets no statements - #225

Open
fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/faithfulness-empty-statements
Open

fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/faithfulness-empty-statements

Conversation

@feiiiiii5

Copy link
Copy Markdown

When the LLM judge returns no statements — or no verdicts for the statements it did return — Faithfulness divided by len(faithfulness) with an empty list, so a single unfaithful-looking answer killed the whole eval run instead of scoring it:

Faithfulness().eval(input="What is the capital of France?", output="", context="Paris is the capital of France.")
# ZeroDivisionError: division by zero  (py/autoevals/ragas.py:998)

The TypeScript scorer already handles this case, in js/ragas.ts:

const score = faithfulness.length
  ? faithfulness.reduce((acc, { verdict }) => acc + verdict, 0) / faithfulness.length
  : 0;

This applies the same guard to the Python async and sync paths, so an empty verdict list scores 0 and the run continues.

Test: pytest py/autoevals/test_ragas.py -k no_statements — the new test fails on 9546b28 with ZeroDivisionError: division by zero at py/autoevals/ragas.py:998 (both the sync and async parametrization) and passes on this branch. The rest of test_ragas.py is unchanged: 9 of its tests fail identically before and after this branch because they call a real OpenAI model and this environment has no API key.

Also ran the repo's configured lint on the two changed files: black --check (clean), ruff check (clean), codespell (clean).

Details

AnswerRelevancy has the same unguarded division at py/autoevals/ragas.py:1152 (/ len(questions)), but the TypeScript twin is unguarded in the same place, so there is no prior art to point at. I left it alone rather than fold a second, differently-evidenced change into this one.

py/autoevals/ragas.py:402 (ContextRelevancy, / len(context)) is likewise unguarded in both languages; that one divides by user input rather than a judge result, so it is a different shape again.

An LLM judge that returns no statements (or no verdicts for them) left
`faithfulness` empty, and the scorer divided by `len(faithfulness)`, so the
whole eval run died with a ZeroDivisionError instead of reporting an
unfaithful answer. The TypeScript scorer already returns 0 in this case.

Guard both the async and sync paths the same way js/ragas.ts does, and cover
them with a regression test.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant