<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Benchmarks on jjshanks.net</title>
    <link>http://www.jjshanks.net/tags/benchmarks/</link>
    <description>Recent content in Benchmarks on jjshanks.net</description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Tue, 28 Jul 2026 09:00:00 -0800</lastBuildDate>
    <atom:link href="http://www.jjshanks.net/tags/benchmarks/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Building a Personal AI Benchmark</title>
      <link>http://www.jjshanks.net/posts/personal-ai-benchmark/</link>
      <pubDate>Tue, 28 Jul 2026 09:00:00 -0800</pubDate>
      <guid>http://www.jjshanks.net/posts/personal-ai-benchmark/</guid>
      <description>&lt;p&gt;&lt;strong&gt;Quick Take&lt;/strong&gt; I stood up a personal agentic AI benchmark using Inspect AI, with tasks running in Docker and OpenRouter as the default inference provider. A first run of six starter tasks against four free models already surprised me: ling-3.0-flash, a model I had never heard of, matched every score while using the fewest tokens.&lt;/p&gt;&#xA;&lt;p&gt;New models ship constantly, and I have no good way of telling whether any of them matter for what I actually do. The published numbers don&amp;rsquo;t help; they&amp;rsquo;re benchmarks I don&amp;rsquo;t run, on tasks I don&amp;rsquo;t have. Meanwhile my own setup runs on auto-pilot: Claude Code with Opus 5 for coding, ChatGPT 5.6 Sol for adversarial reviews, Gemini 3.6 Flash for personal things. It works, but I couldn&amp;rsquo;t defend any of it with evidence.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
