A new benchmark measures how well AI agents can automate economically valuable chores. Human-level AI is still some ways off.
Related Posts
I Watched AI Agents Try to Hack My Vibe-Coded Website
RunSybil, a startup founded by OpenAI’s first security researcher, deploys agents that probe websites for vulnerabilities—part of a new AI era for cybersecurity.
AI’s Hacking Skills Are Approaching an ‘Inflection Point’
AI models are getting so good at finding vulnerabilities that some experts say the tech industry might need to rethink how software is built.
It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
I watched a new tool try to get around the model safeguards of four major frontier companies. You might be surprised by how they performed.

