Benchmarks

    Unified intelligence. Benchmark-leading performance.

    AdaL brings coding, browser use, and long-running agent workflows into one system — with strong results across BU100 + SWE-Pro-Curated-50.

    View coding benchmark source code →
    92% BU1001.09× coding accuracy64% lower cost / success

    BU100 / BU Bench V1

    AdaL Browser Use reaches 92%.

    +2.5pp over the best shown baseline

    BU Bench V1 evaluates browser automation agents on 100 hand-selected tasks. The comparison below mirrors the provided benchmark view and adds AdaL’s current 92% BU100 result.

    92%
    89.5%
    78%
    77%
    74%
    68%
    AdaL Browser Use
    Unified agent
    BrowserCode Opus 4.7
    Open-source agent
    Browser Use Cloud v3 Opus 4.7
    Cloud-hosted agent
    Claude Code + Agent Browser Opus 4.7
    Browser agent setup
    Claude Code + Browser Harness Opus 4.7
    Browser harness setup
    Browser Use Local GPT-5.5
    Local browser agent
    Suite
    100 browser automation tasks
    Score
    Task success rate
    AdaL result
    92% on BU100
    Metric
    AdaL
    Comparison
    Browser Use
    92%
    89.5% best baseline
    Coding accuracy
    50%
    46% Claude Code
    Speed
    209s
    315s Claude Code
    Cost
    $17.10
    $43.09 Claude Code
    Tool calls
    897
    2,323 Claude Code
    Output tokens
    346,969
    617,356 Claude Code
    Cost / success
    $0.68
    $1.87 Claude Code

    Coding / Swe-Bench Pro Curated 50 / Live

    SWE-Pro Coding — Live Leaderboard.

    Latest-run results on the SWE-Bench Pro Curated-50 suite — pass rate, cost, and time from the most recent run per model and thinking-effort configuration. Results update automatically as new runs complete.

    0 live modelslatest run per configuration • updated via CI
    #
    Agent
    Model
    Pass Rate
    Cost
    Time
    Loading live curated-50 leaderboard…

    Methodology: SWE-Bench Pro Curated-50 (50 real-world cases). Infrastructure by Margin Lab. Showing the most recent run per agent × model × thinking effort configuration. Results update automatically via CI. Live from database on page load.

    Coding / Swe-Bench Pro Curated 50 / Trend

    Performance over runs.

    One point per eval run.

    Loading run timeline…

    Grouping and ordering computed server-side from the eval database — one series per agent × model × thinking-effort configuration, most recent 10 runs each.