List from agenticcoding.com
AI for Developers
Tools, techniques and ideas for building software with AI.
How to build a trusted intelligence layer for reliable AI agents
From context engineering to trusted intelligence: a practical architecture for relevant evidence, governed access and human control.

V0.203.4: fix(deps): publish the 0.203 line on agent-core 0.10.2 (0.203.4) (#971)
fix(deps): publish the 0.203 line on agent-core 0.10.2 (0.203.4) v0.203.3 (#970) did not publish. Its packed-consumer check found two agent-interface copies, because agent-core 0.10.3 (2026-10-06) moved to agent-interface 3 in a patch release, and the 0.203 line's ^0.10.2 now resolves it beside the consumer's interface 2.
Evals crash course with Hamel Husain
Watch now (57 mins) | Hamel shows how to use Claude Code and Codex to find real AI failures, build useful evals, and improve your workflows.

The Agents in Production Aren’t Mine. Here’s What Their Server Sees
What our MCP server logs about agents we don’t control, and what my own scheduled agents really costSomewhere between June 3 and September 2, 2026, an agent asked our production MCP server for a tool called GBContent.getItems(sectionId, opts, onOk, onErr). Parentheses, parameter names, callbacks, the whole JavaScript signature, sent as the name of a tool. The server has never had a tool by that name.

V0.203.3: feat(meta-eval): backport the judge gate to 0.203 (0.203.3) (#970)
feat(meta-eval): backport the judge gate to 0.203 (0.203.3) GTM pins agent-runtime 0.291.0 and agent-knowledge 18.0.0, whose agent-eval peer windows end below 0.204. The judge gate shipped in 0.210/0.211 (#961, #963), so GTM's adoption (#1445) broke its peer-floor ship check and was reverted (#1448). No agent-runtime or agent-knowledge release on GTM's interface-2 cohort admits agent-eval 0.211.

Why Your Agent Loop Costs More Than Its Model Price Tag
Opus costs $4 per million input tokens. Haiku costs $1 [1]. Swap an agent from one to the other and you’d expect the bill to drop by roughly 75%. In a tool-heavy agent it often drops by a fraction of that, because the per-token price was never the biggest variable in the equation. The loop around the model is.

Github-v1.2.26: chore: promote v2 to dev
Takes the v2 tree, with packages/console, packages/web, packages/stats, packages/function, infra and sst.config.ts merged up to latest dev, and retargets V2 publishing and deploys from the v2 branch to dev.
Utilitarianism and Autism
This is a crosspost from my substack. Effective altruism has received a lot of media attention since the Department of War CTO of the United States declared that there follow “Americanism, NOT Effective Altruism.” One of the main critiques of effective altruists is to accuse them of being utilitarians. It’s an especially tiring criticism, especially from other academics, who should know better than to spread disinformation. As one of the founders of the movement put it:
Anthropic is cutting off its internal evaluations from the internet
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed "unintended model actions," including submitting a false tip regarding an unsolved murder, that led to the decision.

The Evolution of Fine-Tuning: From Retraining Everything to Rewarding Correctness
Last month I had a model, a dataset, and a GPU already warm. The old reflex said fine-tune. I didn’t. I sat there and asked whether… Continue reading on Towards AI »

Examing Emergent Misalignment in a recurrent LLM with a logit lens
IntroThis post builds on my previous post about mech interp on a recurrent LLM. I found that a logit lens is able to recover a chain-of-thought equivalent from an LLM, loop index is linearly represented, and that steering concepts transfer between loops.
Clips Nightly v0.1.436-0
Auto-updating nightly build of the Clips menu-bar app. The in-app updater fetches its manifest from the clips-nightly-latest release — this versioned release is the immutable source of the signed b
The $80,000 Hallucination: Why RAG Fails at Healthcare Eligibility Verification
Reading comprehension is not transactional verification. When an autonomous intake bot checks a static PDF brochure instead of querying live clearinghouse rails, a routine outpatient surgery becomes an unmitigated financial catastrophe.A 42-year-old patient scheduled an outpatient cervical spine decompression at an ambulatory surgical center.

Why Does AI Code for Problems It Imagines Instead of Yours?
How vague prompts trick AI into solving problems you'll pay for forever

How Can AI Agents Read Untrusted Sources Safely?
Add architectural guardrails around how agents read untrusted sources and use them in the workflow. The post How Can AI Agents Read Untrusted Sources Safely? appeared first on Towards Data Science.
Inheritance of Refusals from Abliterated Models
This project began as part of a short application to Neel's MATS 12 stream. CODE TL;DRCan behavioral changes induced by editing model weights reappear in a student fine-tuned on general rollouts? Inspired by Arthur Conmy's post on hereditary traits, I explored this question using an abliterated Qwen 3.5 9B teacher and smaller students. I find that specific model behaviors are transmitted through distillation, even when those behaviors are induced by techniques based on mechanistic interpretability instead of traditional SFT or prompting.
Release Candidate v1.6.13-rc.98
Automated release candidate build from main.\n\nnpm: npm install gitnexus@rc\nVersion: 1.6.13-rc.98\nTarget base: 1.6.13 (rc #98)\nSource commit (main): a532e08\nRelease commit (versioned tree): 52dcd2c\n\nRelease candidates are pre-stable builds intended for early testing. Stable releases remain on the latest dist-tag. What's Changed 🚨 Security Improve MCP startup compatibility and lazy-load CLI commands by
OAuth 2.0 Token Exchange — Laying the Foundation for AI Agent Authorization
BackgroundIn a previous article, I walked through the three most common OAuth 2.0 grant flows: Authorization Code Flow, PKCE Flow, and Device Flow. We looked at which scenarios and application types each one fits, what problems they solve, and how they differ in implementation.


Showing 20 of 99 articles