docs: replace README plot with hard benchmark v2 (#5752)

Replace the README's BU Bench V1 plot with the supplied Browser Use
Benchmark v2 artwork. Reduce the GPT-6 ASTRA label and 77.3% score text,
retaining the score value.

The caption links to `browser-use/benchmark`, explains that the
benchmark targets the hardest browser tasks, and notes that smaller
models can achieve very high success rates on easier tasks. Preserve the
benchmark repository's qualification that the plotted results cover a
60-task subset. Remove the outdated reference to the plot as an
open-source versus hosted-agent comparison.

Validation: pre-commit checks and `git diff --check` passed; the edited
chart's labels and plotted values were visually checked against the
supplied artwork.


<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Replaces the README's BU Bench V1 plot with the new Browser Use
Benchmark v2 artwork. Updates the caption and surrounding text to
describe the new benchmark, including its focus on the hardest browser
tasks and the 60-task subset used for the plotted results. Also removes
the outdated open-source vs hosted-agent comparison and the "see plot
above" reference in the cloud agent section.

<sup>Written for commit f40fa559fb.
Summary will update on new commits.</sup>

<a
href="https://cubic.dev/pr/browser-use/browser-use/pull/5752?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>

<!-- End of auto-generated description by cubic. -->
This commit is contained in:
Magnus Müller
2026-09-09 11:18:15 -07:00
committed by GitHub
2 changed files with 3 additions and 7 deletions
+3 -7
View File
@@ -129,13 +129,9 @@ Check out the [library docs](https://docs.browser-use.com/open-source/introducti
# Open Source vs Cloud
<picture>
<source media="(prefers-color-scheme: light)" srcset="static/accuracy_by_model_light.png">
<source media="(prefers-color-scheme: dark)" srcset="static/accuracy_by_model_dark.png">
<img alt="BU Bench V1 - LLM Success Rates" src="static/accuracy_by_model_light.png" width="100%">
</picture>
<img alt="Browser Use Benchmark v2 - Mean rubric score by model and cost per task" src="static/hard_benchmark_v2.jpg" width="100%">
We benchmark Browser Use across 100 real-world browser tasks. Full benchmark is open source: **[browser-use/benchmark](https://github.com/browser-use/benchmark)**.
This [very hard benchmark](https://github.com/browser-use/benchmark) targets the hardest browser tasks. On easier tasks, even smaller models can achieve very high success rates. Results shown are from a 60-task subset of BU Bench V2.
Browser Use is also **#1 on the [Odysseys leaderboard](https://odysseysbench.com/leaderboard)** with an 87.4% average, ahead of computer-use agents from OpenAI, Anthropic, Google, and Microsoft. Odysseys measures the agent's performance on 200 long-horizon web tasks.
@@ -145,7 +141,7 @@ Browser Use is also **#1 on the [Odysseys leaderboard](https://odysseysbench.com
- We recommend pairing it with our [cloud browsers](https://docs.browser-use.com/open-source/customize/browser/remote) for leading stealth, proxy rotation, and scaling
**Use the [Fully-Hosted Cloud Agent](https://cloud.browser-use.com?utm_source=github&utm_medium=readme-hosted-agent) (recommended)**
- Much more powerful agent for complex tasks (see plot above)
- Much more powerful agent for complex tasks
- Easiest way to start and scale
- Best stealth with proxy rotation and captcha solving
- 1000+ integrations (Gmail, Slack, Notion, and more)
Binary file not shown.

After

Width:  |  Height:  |  Size: 252 KiB