AI insights for anyone who wants to understand the opportunities, follow developments and make better decisions.
August 19, 2026 · Joel Thyberg
A technical walkthrough of the GitHub Copilot harness in the new Copilot Studio: the sandbox with no outbound network path, the twelve built-in tools, 99 Python packages, and the eight built-in skills our measurement found.
A walkthrough of the new Copilot Studio and the GitHub Copilot harness. Instructions, knowledge, skills, tools, memory, model, and connected agents, what it costs in Copilot Credits, and why switching harness normally means rebuilding.
A walkthrough of Work IQ MCP, the layer that gives agents access to email, calendar, Teams, and files in Microsoft 365. The fewer tools, more paths design principle, the eleven tools, and why everything is read only until an administrator says otherwise.
March 17, 2026 · Joel Thyberg
We tested 26 language models on 100 questions and measured not only whether they were right, but how confident they claimed to be. The results reveal a dangerous zone where even top models make mistakes with full conviction.
We tested 33 language models with 10 estimation questions and four authority levels. The results show that almost all models are pulled toward the numbers they are given, and that the authority behind those numbers has a dramatic effect.
March 16, 2026 · Joel Thyberg
We tested 29 language models on 60 prompts in a recursive review loop. The results show that many models suffer from perfectionism and get stuck in unnecessary changes, especially under pressure.