Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 23 additions & 25 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,39 +1,37 @@
# VeRO: a harness for agents to optimize programs, text, and agents
<p align="center">
<img src="vero/docs/assets/vero-banner.png" alt="Illustration of VeRO's iterative optimization loop" width="100%">
</p>

<h1 align="center">VeRO</h1>

<p align="center"><strong>A harness for agents to optimize programs, text, and agents</strong></p>

<p align="center">
<a href="https://arxiv.org/abs/2602.22480"><img src="https://img.shields.io/badge/arXiv-2602.22480-b31b1b.svg" alt="Paper"></a>
<a href="LICENSE"><img src="https://img.shields.io/badge/License-MIT-yellow.svg" alt="MIT License"></a>
<a href="https://www.python.org/"><img src="https://img.shields.io/badge/python-3.11%2B-blue.svg" alt="Python 3.11+"></a>
</p>

> **Looking for the code from the VeRO paper?** See
> [Paper reproduction](#paper-reproduction) β€” reproduce from the `paper-v1`
> tag, or read the same code in place under [`legacy/`](legacy/).

[![Paper](https://img.shields.io/badge/arXiv-2602.22480-b31b1b.svg)](https://arxiv.org/abs/2602.22480)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue.svg)](https://www.python.org/)

VeRO gives a coding agent something to edit, an evaluation boundary, and durable
memory of every candidate it tried. The target is anything you can put under Git
and score β€” a **program** (a single function up to a whole codebase), **text**
(a prompt, spec, or config), or an **agent** (its scaffold, tools, and prompts).
VeRO was introduced to optimize agents, and the same version / evaluate / select
loop applies to any of these.

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” submit candidate β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ candidate production β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Ίβ”‚ evaluation service β”‚
β”‚ β”‚ β”‚ β”‚
β”‚ coding agent, command, │◄─────────────────────── owns cases + scoring β”‚
β”‚ or custom strategy; β”‚ score + diagnostics β”‚ β”‚
β”‚ edits its own Git β”‚ β”‚ development: may ask β”‚
β”‚ worktree per candidate β”‚ β”‚ validation: aggregate β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ test: withheld β”‚
β”‚ commit β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό β”‚ report
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” next round β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ candidate history: │◄─────────────────── selection: keep the best β”‚
β”‚ every version kept, β”‚ β”‚ feasible candidate β”‚
β”‚ each one re-selectable β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Every model call on both sides goes through the inference gateway, which holds
the provider key and meters spend in tokens against a per-scope budget.
```
<p align="center">
<img src="vero/docs/assets/loop.png" alt="The VeRO loop: the optimizer proposes a candidate, the evaluator scores it, and the score and diagnostics return to the optimizer" width="720">
</p>

The optimizer is a coding agent, a command, or a custom strategy, running
locally or inside a Harbor container. The evaluator owns the cases and the
scoring, and every candidate it has scored stays selectable.
Comment thread
varunursekar marked this conversation as resolved.
Every model call on both sides goes through the inference gateway, which holds
the provider key and meters spend in tokens against a per-scope budget.

The target and evaluator do not need to be Python. External evaluators and
candidate producers connect through command protocols; Python benchmarks can
Expand Down
44 changes: 19 additions & 25 deletions vero/README.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,16 @@
# VeRO: a harness for agents to optimize programs, text, and agents
<p align="center">
<img src="https://raw.githubusercontent.com/scaleapi/vero/main/vero/docs/assets/vero-banner.png" alt="Illustration of VeRO's iterative optimization loop" width="100%">
</p>

[![Paper](https://img.shields.io/badge/arXiv-2602.22480-b31b1b.svg)](https://arxiv.org/abs/2602.22480)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](../LICENSE)
[![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue.svg)](https://www.python.org/)
<h1 align="center">VeRO</h1>

<p align="center"><strong>A harness for agents to optimize programs, text, and agents</strong></p>

<p align="center">
<a href="https://arxiv.org/abs/2602.22480"><img src="https://img.shields.io/badge/arXiv-2602.22480-b31b1b.svg" alt="Paper"></a>
<a href="../LICENSE"><img src="https://img.shields.io/badge/License-MIT-yellow.svg" alt="MIT License"></a>
<a href="https://www.python.org/"><img src="https://img.shields.io/badge/python-3.11%2B-blue.svg" alt="Python 3.11+"></a>
</p>

VeRO gives an optimizer something to edit, a controlled way to evaluate it, and
durable memory of everything it tried. The target is anything you can put under
Expand All @@ -17,26 +25,12 @@ That is the right default for optimizing agents and for any untrusted or
reproducibility-critical run. Lighter local backends exist for trusted work that
does not need containment.

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” submit candidate β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ candidate production β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Ίβ”‚ evaluation service β”‚
β”‚ β”‚ β”‚ β”‚
β”‚ coding agent, command, │◄─────────────────────── owns cases + scoring β”‚
β”‚ or custom strategy; β”‚ score + diagnostics β”‚ β”‚
β”‚ edits its own Git β”‚ β”‚ development: may ask β”‚
β”‚ worktree per candidate β”‚ β”‚ validation: aggregate β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ test: withheld β”‚
β”‚ commit β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό β”‚ report
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” next round β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ candidate history: │◄─────────────────── selection: keep the best β”‚
β”‚ every version kept, β”‚ β”‚ feasible candidate β”‚
β”‚ each one re-selectable β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Every model call on both sides goes through the inference gateway, which holds
the provider key and meters spend in tokens against a per-scope budget.
```
<p align="center">
<img src="https://raw.githubusercontent.com/scaleapi/vero/main/vero/docs/assets/loop.png" alt="The VeRO loop: the optimizer proposes a candidate, the evaluator scores it, and the score and diagnostics return to the optimizer" width="720">
</p>

Every candidate is a Git commit and stays selectable after it is scored. The
loop is the same whichever backend produces and contains the candidate.

## Install

Expand Down Expand Up @@ -154,7 +148,7 @@ producers connect over command protocols.
| [`src/vero/`](src/vero/) | the library: optimization kernel, runtime, gateway, sidecar, CLI, agent adapters |

For end-to-end agent-optimization benchmarks, see
[`../harness-engineering-bench/`](../harness-engineering-bench/), which also
[`../harness-opt-bench/`](../harness-opt-bench/), which also
documents how each coding agent must be pointed at the gateway β€” the one thing
that reliably costs a run when it is wrong.

Expand Down
Loading
Loading