fiper

List from agenticcoding.com

AI for Developers

Tools, techniques and ideas for building software with AI.

Latest stories
  • AI | The Verge Read

    Anthropic is cutting off its internal evaluations from the internet

    After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed "unintended model actions," including submitting a false tip regarding an unsolved murder, that led to the decision.

  • Towards AI - Medium Read

    The Evolution of Fine-Tuning: From Retraining Everything to Rewarding Correctness

    Last month I had a model, a dataset, and a GPU already warm. The old reflex said fine-tune. I didn’t. I sat there and asked whether… Continue reading on Towards AI »

  • LessWrong Read

    Examing Emergent Misalignment in a recurrent LLM with a logit lens

    IntroThis post builds on my previous post about mech interp on a recurrent LLM. I found that a logit lens is able to recover a chain-of-thought equivalent from an LLM, loop index is linearly represented, and that steering concepts transfer between loops.

  • Release notes from agent-native Read

    Clips Nightly v0.1.436-0

    Auto-updating nightly build of the Clips menu-bar app. The in-app updater fetches its manifest from the clips-nightly-latest release — this versioned release is the immutable source of the signed b

  • Towards AI - Medium Read

    The $80,000 Hallucination: Why RAG Fails at Healthcare Eligibility Verification

    Reading comprehension is not transactional verification. When an autonomous intake bot checks a static PDF brochure instead of querying live clearinghouse rails, a routine outpatient surgery becomes an unmitigated financial catastrophe.A 42-year-old patient scheduled an outpatient cervical spine decompression at an ambulatory surgical center.

  • Claude Code Conversations Read

    Why Does AI Code for Problems It Imagines Instead of Yours?

    How vague prompts trick AI into solving problems you'll pay for forever

  • Towards Data Science Read

    How Can AI Agents Read Untrusted Sources Safely?

    Add architectural guardrails around how agents read untrusted sources and use them in the workflow. The post How Can AI Agents Read Untrusted Sources Safely? appeared first on Towards Data Science.

  • LessWrong Read

    Inheritance of Refusals from Abliterated Models

    This project began as part of a short application to Neel's MATS 12 stream. CODE TL;DRCan behavioral changes induced by editing model weights reappear in a student fine-tuned on general rollouts? Inspired by Arthur Conmy's post on hereditary traits, I explored this question using an abliterated Qwen 3.5 9B teacher and smaller students. I find that specific model behaviors are transmitted through distillation, even when those behaviors are induced by techniques based on mechanistic interpretability instead of traditional SFT or prompting.

  • Release notes from GitNexus Read

    Release Candidate v1.6.13-rc.98

    Automated release candidate build from main.\n\nnpm: npm install gitnexus@rc\nVersion: 1.6.13-rc.98\nTarget base: 1.6.13 (rc #98)\nSource commit (main): a532e08\nRelease commit (versioned tree): 52dcd2c\n\nRelease candidates are pre-stable builds intended for early testing. Stable releases remain on the latest dist-tag. What's Changed 🚨 Security Improve MCP startup compatibility and lazy-load CLI commands by

  • Towards AI - Medium Read

    OAuth 2.0 Token Exchange — Laying the Foundation for AI Agent Authorization

    BackgroundIn a previous article, I walked through the three most common OAuth 2.0 grant flows: Authorization Code Flow, PKCE Flow, and Device Flow. We looked at which scenarios and application types each one fits, what problems they solve, and how they differ in implementation.

  • Discover AI [YouTube]Video Read

    AI Agent Self-Learning per 1000 US$

  • ClaudeAI [Reddit] Read

    I made my Claude Code sub-agents a video call. Weirdly, it’s the best way I’ve found to follow what they’re doing

    I built this with Claude Code. It’s a free and open source VS Code extension that shows a Claude Code session as a live video call, so I can see what each sub-agent is doing. It also work with other editors or standalone with claude app as a node server as well.

  • Release notes from jcode Read

    Release v0.94.0 · 1jehuang/jcode

    Queued commands Highlights Ctrl+Enter queues slash and ! commands to run after the current turn, and /clear drops anything still queued Alt+G opens the project's git commit log, and diff cycling moves to /diff Harness API clients can observe and cancel messages sent mid-turn Improvements The startup frame settles immediately instead of popping in over about 200ms, with session facts shown on the first frame The Overview widget stays visible after startup

  • LangChain [Reddit] Read

    I independently audited a RAG benchmark. 15 scoring mismatches revealed a flaw in its main comparison — the maintainer confirmed and fixed it

    ​ I recently completed an independent retrieval audit of RouteMind, an open-source project exploring document routing as an alternative to traditional RAG retrieval. The interesting part wasn't finding a dramatic regression or proving that one architecture was better. It was discovering that two approaches were being evaluated under different definitions of success. Here's what happened. The project had 700 evaluation questions comparing traditional RAG, RAG + reranking, and document-routing approaches.

  • ClaudeAI [Reddit] Read

    I'm a 9th grader and I used Claude to prove a geometry conjecture: the rhombicosidodecahedron can't pass through a copy of itself

    Last year, mathematicians found the first convex shape that can't be pushed through a hole in a copy of itself (the "Noperthedron"). In 2021 the same researchers conjectured that the rhombicosidodecahedron, a classic Archimedean solid, also can't do it, but nobody had proved it.

  • Agentbuild.ai Read

    Complexity Failure: The Hidden Cost of Evaluation Debt in Multi-Agent Systems

    When a multi-agent decision goes wrong, most teams cannot prove which agent caused it. Every agent you add makes that harder. What do you do?

  • Towards AI - Medium Read

    The Summer AI Agents Started Knocking on Government Doors

    OpenAI’s own agents wandered i⁠nto US and Aus​tralian government‍ websites, b⁠orrowi​ng keys stra​ngers had left​ lying around online… Continue reading on Towards AI »

  • AI | The Verge Read

    AI agent makers are promising privacy — will they deliver?

    At this year's OpenAI DevDay, CEO Sam Altman unveiled the company's new AI agent Dots - and told the crowd that the company wants to "set a new standard for privacy in frontier AI." OpenAI would spend the day taking veiled shots at Meta's Muse, its primary competitor, for failing to keep users' data safe. Yet Muse itself, a couple of months earlier, had launched as a supposedly safer alternative to predecessor OpenClaw - with CEO Mark Zuckerberg promising it was "built from the ground up for privacy and security."

  • The New Stack Read

    Microsoft skipped OpenAI’s decision model and built its own on Alibaba’s Qwen

    Microsoft’s Decision-1 matches Jev’s price and skips OpenAI’s model. Here’s who it’s built for and what it still leaves out for developers building agents.

  • Release notes from prime-agent Read

    Nightly build v0.10.1-beta.8

    Built from the continuous train at parent 2ee7e62 (the newest successful continuous build on main) via the nightly workflow's tag-only commit. The -beta* tag publishes the nightly channel (beta.json) through the release.yml promote gate.

Load more stories

Showing 20 of 100 articles