Curated developer articles, tutorials, and guides — auto-updated hourly

New benchmark shows frontier LLMs can only recover research ideas from bibliographies 3-15% of the t...

A new benchmark shows frontier LLMs recover research paper ideas from bibliographies alone at just 3...


Introduction Code duplication detection is a cornerstone of modern software development,...


Open-weight models now match closed frontiers on independent benchmarks. The question is no longer '...


The top ARC-AGI-3 entry on ARC Prize's public leaderboard is an NVIDIA-labelled agent scoring 85.1% ...


LongRCA Bench collects 1,140 genuinely failed agent runs averaging 145 steps each, with human labels...


A new benchmark scores AI video generation on two separate axes and finds the best model reaches 91....


A new benchmark of 119 real scientific software tasks keeps its grading tests private, and the top a...


Microsoft researchers built a system that reads an agent's failure traces, writes structured patches...