Benchmarks
Unified intelligence. Benchmark-leading performance.
AdaL brings coding, browser use, and long-running agent workflows into one system — with strong results across BU100 + SWE-Pro-Curated-50.
View coding benchmark source code →BU100 / BU Bench V1
AdaL Browser Use reaches 92%.
BU Bench V1 evaluates browser automation agents on 100 hand-selected tasks. The comparison below mirrors the provided benchmark view and adds AdaL’s current 92% BU100 result.
Coding / Swe-Bench Pro Curated 50 / Live
SWE-Pro Coding — Live Leaderboard.
Latest-run results on the SWE-Bench Pro Curated-50 suite — pass rate, cost, and time from the most recent run per model and thinking-effort configuration. Results update automatically as new runs complete.
Methodology: SWE-Bench Pro Curated-50 (50 real-world cases). Infrastructure by Margin Lab. Showing the most recent run per agent × model × thinking effort configuration. Results update automatically via CI. Live from database on page load.
Coding / Swe-Bench Pro Curated 50 / Trend
Performance over runs.
One point per eval run.
Grouping and ordering computed server-side from the eval database — one series per agent × model × thinking-effort configuration, most recent 10 runs each.