Kernel

february

NewsCatherine Jue

Evaluating computer use models with Anthropic

Kernel helped Anthropic benchmark Sonnet 4.6, showing how fast, reliable browser infrastructure is used to evaluate computer use models.

Evaluating computer use models with Anthropic

Today, Anthropic is rolling out Sonnet 4.6, a model release that sets a new bar for what computer use models can do. Evaluating that kind of model at scale means throwing it at real, messy websites, which is why Anthropic relied on Kernel to put Sonnet 4.6 to the test.

Finding the toughest login pages on the internet

We recently released Managed Auth, a standardized way to let agents log in and stay logged in across the internet. Part of it includes an agent whose only job is to hunt down a site’s login page using our browser infrastructure. Success is binary: did the agent reach the login page or not. That made it a perfect real-world stress test for Sonnet 4.6.

So we built a focused “find the login page” eval with 254 different sites, including some of the most painful logins we’ve encountered on the internet. We ran this eval on multiple Anthropic models and measured how often each one actually got to the right place. Sonnet 4.6 hit the login page 79.1% of the time on this benchmark, and stood out as the most accurate.

Available today

This is the first step in our work with Anthropic on evaluating computer use models. As these models get better at navigating the internet on your behalf, they need fast, reliable browser infrastructure to do it. You can see how this works by using Sonnet 4.6 on Kernel today.

more blog posts

view all