2 updates from Claude and Anthropic in January 2026, from product launches to research papers, each with a short satori. note on what it means for a European business.

Anthropic redesigned its take-home test three times as Claude improved, so that interviews still say something about human ability.
Tech companies need to rethink their coding interviews, because this is no longer theory. AI handles take-homes, so the real question is what we are testing now.

Evaluating agents requires a combination of code graders, model graders and human review, for both results and interaction quality.
Evaluating agents is where we get stuck with many clients. This is a good base map. Save the link for the next proof of concept.
We use analytics cookies to see which pages people read, and marketing cookies stay off unless you tick them under Details.