Skip to content

Stop HTML-escaping scorer values in the ragas prompts - #227

Open
fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/ragas-prompts-no-html-escaping
Open

fei (feiiiiii5) wants to merge 1 commit into
braintrustdata:mainfrom
feiiiiii5:fix/ragas-prompts-no-html-escaping

Conversation

@feiiiiii5

Copy link
Copy Markdown

The ragas scorers HTML-escape the values they interpolate, so the judge reads a corrupted prompt:

fake.ragas.ContextEntityRecall({ text: 'Is 3 < 5 & 2 > 1? "q"' })
// reaches the model as: text: Is 3 &lt; 5 &amp; 2 &gt; 1? &quot;q&quot;

js/render-messages.ts renders values verbatim and pins that with a test — "should never HTML-escape values, regardless of mustache syntax" — but js/ragas.ts still calls mustache.render with the default escaper at all eight prompt sites. So the same value reaches the judge untouched through the LLMClassifier scorers (Factuality, ClosedQA, …) and entity-encoded through the ragas ones. For anything containing <, >, &, ", ' or / — scraped HTML, source code, maths — the prompt no longer matches the input, and a recorded trace no longer reproduces what was evaluated. Tag names read as literal text can also shift entity extraction.

The fix exports the escapeValue helper that already exists in render-messages.ts and passes it at the eight call sites, so both paths render the same way.

Test: pnpm run test --run js/ragas.test.ts -t "Ragas prompt rendering" fails on 9546b28 for the three scorers above and passes here (4 passed / 6 skipped); the new tests capture the outgoing user message and assert on it. Suite: 69 passed before, 73 after, with the same 8 pre-existing failures either way — all of them Missing credentials … set the OPENAI_API_KEY, which CI supplies from secrets. pnpm run build (tsup, CJS + ESM + DTS) passes, and pre-commit on the three changed files passes. There is no lint or typecheck script in this repo; npx tsc --noEmit reports one pre-existing error in js/render-messages.test.ts that this change does not touch.

Two things to push back on if you disagree:

  • This makes JS differ from the Python implementation. py/autoevals/ragas.py renders with chevron, which HTML-escapes by default (verified on chevron 0.14.0, the pinned version), so today both languages escape and the bug is symmetric. The argument for changing JS is the repository's own test and the escape objects to stringified in autoevals #110 commit that fixed render-messages, not cross-language parity. If you would rather have parity, the alternative is to bypass chevron's escaping in Python ({{&text}} or an escape hook) — happy to do that as a follow-up.
  • {{statements}} changes shape. It is a string[], so with the verbatim escaper it renders as JSON instead of comma-joined. That moves ragas towards Python (which renders the list directly), but it is a real prompt change at that one site; say the word and I will special-case it.

`js/render-messages.ts` renders values verbatim and pins that with a test
("should never HTML-escape values, regardless of mustache syntax"), but
`js/ragas.ts` still calls `mustache.render` with the default escaper at all
eight prompt sites. The same value therefore reaches the judge model
untouched through the LLMClassifier scorers and entity-encoded through the
ragas ones:

    text: Is 3 &lt; 5 &amp; 2 &gt; 1? &quot;q&quot;

For anything scraped from HTML, code or maths that is a corrupted prompt, and
the recorded trace no longer reproduces the input it came from.

Export the existing `escapeValue` and pass it at the eight call sites, which
is what render-messages already does.

One behaviour change beyond escaping: `{{statements}}` is a `string[]`, so
with the verbatim escaper it renders as JSON rather than comma-joined. That
moves ragas towards the Python implementation, which renders the list
directly; if you would rather keep the old comma-joined form, that is the one
site to special-case.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant