Exploring a new AI paper every week.

Dr. Robert Li and Saadia Carbis trace cutting-edge AI research back to the real world.

Skill Issue

What are agent skills good for? What are the best (and worst) ways to use them? We explore Demystifying Agent Skills: Why They Work-Until They Don’t (Jiang, et al., 2026) to find out.

https://doi.org/10.48550/arXiv.2608.14036

Season 1 · Episode 20 · August 20, 2026

An Open-Source Heavyweight

Is it practical to run the most efficient 3T class models? What are the implications of having 3T class model that’s open-source? We explore Kimi K3: Open Frontier Intelligence (Kimi Team, 2026) to find out.

https://doi.org/10.48550/arXiv.2607.24653

Season 1 · Episode 19 · August 6, 2026

Claude is Cool, GPT Likes to Gossip

Can AI write fiction like a human? We explore StoryScope: Investigating idiosyncrasies in AI fiction (Russell, et al., 2026) to find out.

https://doi.org/10.48550/arXiv.2604.03136

Season 1 · Episode 18 · July 28, 2026

The Goldblum Effect

Do Anthropic’s models contain a global workspace? If so, what are the implications? We explore Verbalizable Representations Form a Global Workspace in Language Models (Gurnee, et al., 2026) to find out.

Verbalizable Representations Form a Global Workspace in Language Models

Season 1 · Episode 17 · July 10, 2026

Speculative Intelligence

What are the pathways from here to AGI and ASI? Can anyone agree on a definition for AGI? We explore From AGI to ASI (Genewein, et al., 2026) to find out.

https://doi.org/10.48550/arXiv.2606.12683

Season 1 · Episode 16 · June 28, 2026

Disentangling Model From Harness

How can a self-evolving harness benefit a model’s capabilities? We explore Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents (Lin, et al., 2026) to find out.

https://doi.org/10.48550/arXiv.2605.30621

Season 1 · Episode 15 · May 28, 2026

Superhuman Step Change

Is Mythos all its hyped up to be? We explore System Card: Claude Mythos Preview (Anthropic, 2026) to find out.

System Card: Claude Mythos Preview

Season 1 · Episode 14 · May 28, 2026

Opening the Black Box

We’re joined by Gareth O’Shea, to ask:

Could it be possible to run an MRI on an LLM while it’s thinking? We explore Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models (Qwen Team, 2026) to find out.

https://doi.org/10.48550/arXiv.2605.11887

Season 1 · Episode 13 · May 21, 2026

Talmudic Post-Training

Should we align models based on human preference or human behaviour? We explore Alignment Makes Language Models Normative, Not Descriptive (Shapira, et al., 2026) to find out.

https://doi.org/10.48550/arXiv.2603.17218

Season 1 · Episode 12 · March 26, 2026

Specialisation and Art

How does extra domain specific pre-training contribute to code specific models? We explore Qwen3-Coder-Next Technical Report (Qwen Team, 2026) to find out.

https://doi.org/10.48550/arXiv.2603.00729

Season 1 · Episode 11 · March 17, 2026