<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet type="text/xsl" href="/feed.xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Arbiter Machinae · Research notes</title>
  <link href="https://kiwimaddog2020.github.io/feed.xml" rel="self"/>
  <link href="https://kiwimaddog2020.github.io/"/>
  <id>https://kiwimaddog2020.github.io/</id>
  <updated>2026-06-15T04:41:40Z</updated>
  <author><name>Kevin Madson</name></author>
  <entry>
    <title>An honest eval harness, and a controlled experiment on what makes AI refinement actually improve.</title>
    <link href="https://kiwimaddog2020.github.io/trutina/"/>
    <id>https://kiwimaddog2020.github.io/trutina/</id>
    <updated>2026-06-15T04:41:40Z</updated>
    <summary>An honest eval harness, and a controlled experiment on what makes AI refinement actually improve.</summary>
  </entry>
  <entry>
    <title>Do AI code reviewers have complementary blind spots?</title>
    <link href="https://kiwimaddog2020.github.io/decorrelation-study/"/>
    <id>https://kiwimaddog2020.github.io/decorrelation-study/</id>
    <updated>2026-06-13T23:56:27Z</updated>
    <summary>Do AI code reviewers have complementary blind spots?</summary>
  </entry>
  <entry>
    <title>How to know if your AI is actually working</title>
    <link href="https://kiwimaddog2020.github.io/evaluating-claude/"/>
    <id>https://kiwimaddog2020.github.io/evaluating-claude/</id>
    <updated>2026-06-13T07:27:28Z</updated>
    <summary>How to know if your AI is actually working</summary>
  </entry>
  <entry>
    <title>The gate that edits a gate is rejected by the gate it edits</title>
    <link href="https://kiwimaddog2020.github.io/self-improvement-gate/"/>
    <id>https://kiwimaddog2020.github.io/self-improvement-gate/</id>
    <updated>2026-06-13T06:59:36Z</updated>
    <summary>The gate that edits a gate is rejected by the gate it edits</summary>
  </entry>
  <entry>
    <title>Ten rounds of an agent improving one game</title>
    <link href="https://kiwimaddog2020.github.io/oneshot-bench/"/>
    <id>https://kiwimaddog2020.github.io/oneshot-bench/</id>
    <updated>2026-06-13T06:48:28Z</updated>
    <summary>Ten rounds of an agent improving one game</summary>
  </entry>
  <entry>
    <title>Rating your own work without lying to yourself</title>
    <link href="https://kiwimaddog2020.github.io/craft-fit-rating/"/>
    <id>https://kiwimaddog2020.github.io/craft-fit-rating/</id>
    <updated>2026-06-13T06:19:19Z</updated>
    <summary>Rating your own work without lying to yourself</summary>
  </entry>
  <entry>
    <title>An eval framework where the doer never grades itself</title>
    <link href="https://kiwimaddog2020.github.io/trust-weighted-evals/"/>
    <id>https://kiwimaddog2020.github.io/trust-weighted-evals/</id>
    <updated>2026-06-12T22:30:01Z</updated>
    <summary>An eval framework where the doer never grades itself</summary>
  </entry>
  <entry>
    <title>Audioscan: one decode pass instead of three ffmpeg shellouts</title>
    <link href="https://kiwimaddog2020.github.io/audioscan/"/>
    <id>https://kiwimaddog2020.github.io/audioscan/</id>
    <updated>2026-06-03T06:25:29Z</updated>
    <summary>Audioscan: one decode pass instead of three ffmpeg shellouts</summary>
  </entry>
</feed>
