<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Fabio Nonato de Paula — Blog</title>
  <subtitle>Thoughts on AI/ML infrastructure, developer tools, cybersecurity, and open source.</subtitle>
  <link href="https://nonatofabio.github.io/feed.xml" rel="self"/>
  <link href="https://nonatofabio.github.io/blog/" rel="alternate"/>
  <id>https://nonatofabio.github.io/</id>
  <updated>2026-09-21T00:00:00Z</updated>
  <author>
    <name>Fabio Nonato de Paula</name>
    <uri>https://nonatofabio.github.io/</uri>
  </author>
  <entry>
    <title>A Fruit Fly Brain Plays Doom</title>
    <link href="https://nonatofabio.github.io/blog/posts/doomfly_autonomous.html" rel="alternate"/>
    <id>https://nonatofabio.github.io/blog/posts/doomfly_autonomous.html</id>
    <published>2026-09-21T00:00:00Z</published>
    <updated>2026-09-21T00:00:00Z</updated>
    <summary>A network wired like a fruit fly's brain plays five Doom scenarios. The Strands harness built it, ran the GPU fleet, and reported two negative results. My part of the transcript is a dozen sentences.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="strands"/>
    <category term="ml"/>
    <category term="reinforcement-learning"/>
    <content type="html">&lt;p&gt;Imagine I opened this post by telling you I got a fly to play Doom. Would you believe me?&lt;/p&gt;
&lt;figure&gt;
  &lt;video src=&quot;../../assets/doomfly/doomfly_reel.mp4&quot;
         poster=&quot;../../assets/doomfly/doomfly_reel_poster.jpg&quot;
         autoplay muted loop playsinline preload=&quot;metadata&quot;
         data-missing=&quot;Reel not available yet.&quot;&gt;
  &lt;/video&gt;
  &lt;figcaption&gt;The five ViZDoom scenarios, one clip each. Every clip is the median episode of ten, so this is typical play and not a highlight reel.&lt;/figcaption&gt;
&lt;/figure&gt;&lt;p&gt;You should not. The thing playing up there is not a fly, but its wiring diagram comes from one. It is a recurrent network built on a real fruit fly connectome: 49,393 neurons and 9 million synapses, all frozen, with one learned gain per synapse. An ordinary convolutional network feeds it the game frames and an ordinary MLP turns its activity into moves, and those two hold more than half of the parameters. So read every result in this post as &amp;quot;a network constrained to the fly connectome&amp;quot; and never as &amp;quot;a fly&amp;quot;. Yes, the title breaks that rule. That is the hook, and the rest of the post is the correction.&lt;/p&gt;
&lt;p&gt;Now, here is the thing: I did not write the code, the infrastructure or the first drafts of the docs. All of it came out of the &lt;a href=&quot;https://github.com/strands-agents/harness-sdk&quot;&gt;Strands harness&lt;/a&gt;, an open-source coding agent from the Strands Agents team at AWS that you run from the terminal. It keeps its own state on disk between sessions and runs tools, tests and cloud deployments on its own. Here it wrote the pipeline that turns the fly&amp;#39;s wiring into a network and the five ordinary Doom agents that served as teachers. It wrote the trainer that taught the fly-wired network to copy them, and the second training stage that tried to push it further. It also wrote the cloud setup that ran a GPU fleet in three AWS regions, and it recorded the footage. Five days, 28 commits, about 3,300 lines of Python and shell, twelve fine-tuning runs and three control runs.&lt;/p&gt;
&lt;h2&gt;What I actually typed&lt;/h2&gt;
&lt;p&gt;The harness keeps its own session state on disk, so I went back and read my side of the transcript. It is short. It started with one prompt, typos and all:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I just saw this: &lt;a href=&quot;https://huggingface.co/mlabonne/chessfly&quot;&gt;https://huggingface.co/mlabonne/chessfly&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I want to make the same thing but to play doom! Look at the skills and memory and have a set up ready for us to train it using my fnp3 aws accountl. Also create a infographic style tutorial on how it works, both the conectome model as well as the training. Use this current directory as a infra/cdk for training and training source and data ETL pipeline source. Once we complete all, we should have this is a private github repo. If you find a way to start training while I&amp;#39;m away, go for it!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Across the seven sessions that followed, this is most of what I said:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Go for it!&lt;/p&gt;
&lt;p&gt;Can you check how are we doing on the training of doomfly?&lt;/p&gt;
&lt;p&gt;Do you have suggestions to speed up the training process?&lt;/p&gt;
&lt;p&gt;how do I use tb to see all the runs so far? &lt;em&gt;(tb is TensorBoard)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Let&amp;#39;s stop and revert all this we should not pursue this anymore&lt;/p&gt;
&lt;p&gt;Sorry, wrong session!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;quot;Let&amp;#39;s stop and revert all this&amp;quot; was me killing an easter egg I had asked for an hour earlier: a Doom mod with a fly paw for the player&amp;#39;s hand and Stan the frog, the Strands mascot, as the enemies. The harness reverted the whole thing, previews and tests included, and went back to the training run. &amp;quot;Sorry, wrong session!&amp;quot; came right after I pasted a request meant for a different agent, about a different project.&lt;/p&gt;
&lt;p&gt;I ran all of it from the terminal with the &lt;code&gt;strands&lt;/code&gt; command, and I changed the model underneath it as I went. The first sessions ran on Claude Fable 5.1 on Amazon Bedrock. By the last day the same sessions were running on GPT-6 Astra at the highest reasoning effort. The transcript does not care which model is behind it, and neither does the repo.&lt;/p&gt;
&lt;h2&gt;Between my sentences&lt;/h2&gt;
&lt;p&gt;Everything else in the transcript is the harness at work, and reading it back showed me how much the Strands harness worked for me.&lt;/p&gt;
&lt;p&gt;Each session persisted to disk, and the next one opened with a summary of the last. Seven sessions over five days read like one conversation, and I could close the laptop at night and pick up in the morning where it left off. Long tool outputs were paged out of the conversation as it grew and fetched back when the agent asked for them. About two hundred of those are still sitting on disk, and without that mechanism the logs from day one alone would have ended the project.&lt;/p&gt;
&lt;p&gt;Background tasks paid off on the last day. While a two-hour evaluation ran on my Mac, the same agent widened the model&amp;#39;s output from 22 Doom actions to 23 for the next experiment and wrote nine tests for it. It also confirmed that the widened model played the old scenarios bit-for-bit the same as before. The repo has an &lt;code&gt;AGENTS.md&lt;/code&gt; that says never push, launch, terminate or delete without being told, and the harness asked every time. &amp;quot;Go for it!&amp;quot; was me answering. It also loaded my own &lt;a href=&quot;https://github.com/nonatofabio/claude-writing-skills&quot;&gt;writing and verification skills&lt;/a&gt; and the MCP tools I already use, which is why the docs it produced read like mine.&lt;/p&gt;
&lt;p&gt;None of that is exotic on its own. What I had not seen before is all of it in one command, with nothing to wire up. The same agent is two lines of Python if you want it inside a script, and after this week I would not think twice about starting a project that way.&lt;/p&gt;
&lt;h2&gt;The part I trust&lt;/h2&gt;
&lt;p&gt;An agent that returns a plausible positive result is hard to trust. This one returned two negative results, with the run prefixes and standard deviations attached, because its rules say that negative results go in &lt;code&gt;docs/&lt;/code&gt; and not in the bin.&lt;/p&gt;
&lt;p&gt;A short map of what the harness did, so the two results make sense. It first trained five ordinary Doom agents with PPO, a standard reinforcement learning method, one per scenario. Those are the teachers. It then distilled the teachers into one fly-wired network, the student, by having the student imitate their moves. The student ended up playing about as well as its teachers. The question the project was built for was whether a second stage of reinforcement learning, GRPO, could push it past that.&lt;/p&gt;
&lt;p&gt;First, it could not. Twelve GRPO runs, one setting changed per run, and not one beat the student by more than the noise in the evaluation. GRPO keeps the new policy close to the student with a penalty. The penalty strength that stopped the policy from wandering into worse play was the same strength that stopped it from improving at all.&lt;/p&gt;
&lt;p&gt;Second, and this one stings, the connectome does nothing measurable. The harness suggested and built two controls. One keeps the same graph but shuffles its edges, so every neuron keeps its degree and its sign and the biology is gone. The other removes the neuron layer entirely.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Backbone&lt;/th&gt;
&lt;th&gt;Parameters&lt;/th&gt;
&lt;th&gt;Training steps per second&lt;/th&gt;
&lt;th&gt;Doom scores&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;Real connectome&lt;/td&gt;
&lt;td&gt;48.2 M&lt;/td&gt;
&lt;td&gt;646&lt;/td&gt;
&lt;td&gt;baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shuffled edges, two seeds&lt;/td&gt;
&lt;td&gt;48.2 M&lt;/td&gt;
&lt;td&gt;1,454&lt;/td&gt;
&lt;td&gt;same, within noise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No connectome at all&lt;/td&gt;
&lt;td&gt;27.7 M&lt;/td&gt;
&lt;td&gt;14,686&lt;/td&gt;
&lt;td&gt;same, within noise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;One scenario as an example, ten episodes each: on &lt;code&gt;defend_the_center&lt;/code&gt; the real connectome scored 18.6 ± 1.8, the two shuffled graphs 19.0 and 18.4, and the model with no connectome 18.3. The other four scenarios look the same. The fly wiring makes training 22 times slower and buys nothing on these five scenarios. That is why the caption says &amp;quot;wired like a fruit fly&amp;#39;s brain&amp;quot; and not &amp;quot;a fruit fly&amp;#39;s brain&amp;quot;. The harness wrote that vocabulary rule into &lt;code&gt;AGENTS.md&lt;/code&gt; on day two, before either result existed. It would have been a hard rule to write after a result I disliked.&lt;/p&gt;
&lt;h2&gt;What&amp;#39;s next&lt;/h2&gt;
&lt;p&gt;Scores on the five scenarios stopped moving, and three of them sit at their ceilings, so the next test is a real level. The student already plays Freedoom II MAP01 zero-shot. With its aiming prior it survives twice as long as random and picks up items. It still dies in eight of ten episodes and never leaves the first two rooms. The wider output and a slot for a sixth scenario are done and tested. Next comes GRPO on that map with all three backbones: real, shuffled and none. If the wiring is ever going to matter, it will be where the policy has to learn rather than imitate. If the network learns anything there, the plan is fly versus fly in a duel.&lt;/p&gt;
&lt;p&gt;The code is public at &lt;a href=&quot;https://github.com/nonatofabio/doomfly-rl&quot;&gt;nonatofabio/doomfly-rl&lt;/a&gt;, and the trained checkpoints are on &lt;a href=&quot;https://huggingface.co/fabiononato/doomfly-rl&quot;&gt;Hugging Face&lt;/a&gt; with the connectome file they need. If you want to watch a network shaped like a fly play Doom, start there.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Thanks for reading.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Keep it Awesome!&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>AI Generated Pixel Art Needs a Build System, Not Better Prompts</title>
    <link href="https://nonatofabio.github.io/blog/posts/pixel_art_build_system.html" rel="alternate"/>
    <id>https://nonatofabio.github.io/blog/posts/pixel_art_build_system.html</id>
    <published>2026-09-07T00:00:00Z</published>
    <updated>2026-09-07T00:00:00Z</updated>
    <summary>I rebuilt my Pygame shooter in Godot 4 with an agent. The 95k lines of code were the easy part. Keeping four pictures of the same orc consistent was not.</summary>
    <category term="ai"/>
    <category term="gamedev"/>
    <category term="godot"/>
    <category term="agents"/>
    <category term="pipelines"/>
    <content type="html">&lt;p&gt;I rebuilt my tiny homage to Warhammer 40k this week. The original, Hive City Rampage, was a Python/Pygame monstrosity, a top-down grimdark shooter that ran but was a prototype. Now there&amp;#39;s a sequel, Ashgate Siege, built in Godot 4: gothic isometric, two missions, Mac and Android.&lt;/p&gt;
&lt;p&gt;I have a confession though. The 95k lines of GDScript, 8 commits, roughly five hours, was written by an AI agent under my direction. The sprites are generated too. I&amp;#39;m disclaiming that up front because the interesting part is what came out of it.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;../../assets/hcr2/title.jpg&quot; alt=&quot;Hive City Rampage II title screen&quot;&gt;&lt;/p&gt;
&lt;h2&gt;The easy part was the code&lt;/h2&gt;
&lt;p&gt;This is the simple finding. An engine port is exactly the shape of problem my agents are good at: the target is well documented, the semantics known, and correctness is checkable by running the thing. Godot 4 plus GDScript is heavily represented in training data for any model. Zero agent struggle.&lt;/p&gt;
&lt;h2&gt;The hard part was four pictures of the same orc&lt;/h2&gt;
&lt;p&gt;Here&amp;#39;s the hill I&amp;#39;m willing to die on: for AI-assisted games, 2D sprite games are harder than 3D.&lt;/p&gt;
&lt;p&gt;That sounds backwards, so: in 3D, the engine guarantees coherence. One mesh, one material, one light rig, and every frame of animation is consistent because it&amp;#39;s the same object being transformed. The renderer is doing the work.&lt;/p&gt;
&lt;p&gt;In sprite based games there is no shared object. Each sprite is an independent generated map of pixels. So:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Frame 2 of a walk cycle can be a different character than frame 1.&lt;/li&gt;
&lt;li&gt;The light source can move between sprites in the same scene.&lt;/li&gt;
&lt;li&gt;Proportions drift. Your massive size boss shrinks.&lt;/li&gt;
&lt;li&gt;Limbs get cropped at cell edges.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;None of that is caught by anything automatically. It just ships, and the game looks odd, like a badly executed collage.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;../../assets/hcr2/ashgate.jpg&quot; alt=&quot;Ashgate Siege gameplay: a firefight around a burning signal relay&quot;&gt;&lt;/p&gt;
&lt;p&gt;That screenshot has maybe a dozen generated sprites in it at once. Every one of them came out of a separate generation, and the only reason they read as one scene is the pipeline below.&lt;/p&gt;
&lt;h2&gt;Prompts as interface specs&lt;/h2&gt;
&lt;p&gt;The fix was to stop treating generation as commissioning art and start treating it as calling an API with a strict schema. Every prompt pins the geometry:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Production game sprite sheet: exactly 4 columns x 3 rows, 1536x1024 canvas,
equal cells, solid pure magenta #FF00FF background for color-key import.
NO text, shadows, smoke, or border. Every figure entirely within its own cell
with ample margin; feet centered at same baseline in each row. All figures face
screen RIGHT in three-quarter isometric view [...] Four columns are four
coherent walking-cycle poses: left foot forward, passing, right foot forward,
passing.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Each sentence has to do some work:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Magenta #FF00FF, not transparency.&lt;/strong&gt; Alpha comes back unreliable. A color key is deterministic to strip.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explicit grid and canvas.&lt;/strong&gt; The importer crops on fixed coordinates. If the grid drifts, every sprite is off.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;quot;feet centered at same baseline in each row.&amp;quot;&lt;/strong&gt; Without it the character bobs while walking.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Named poses.&lt;/strong&gt; &amp;quot;Walking animation&amp;quot; returns four unrelated drawings. Naming the four phases is what makes them a cycle.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Negative constraints.&lt;/strong&gt; &amp;quot;No text, shadows, smoke, or border&amp;quot; because models add decoration that breaks the silhouette.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The importer is a compiler, and tests are CI&lt;/h2&gt;
&lt;p&gt;If the prompt is a spec, something has to enforce it. Generated atlases go into &lt;code&gt;art/&lt;/code&gt;, an importer crops and scales them into runtime assets, and a native test suite checks the result: 2,972 reference and combat assertions, 17,112 seam combinations, plus a 600-frame combat simulation. &lt;code&gt;make test&lt;/code&gt; runs it.&lt;/p&gt;
&lt;p&gt;The seam check is the one I lean on. It takes the belt pixels from the leg sprite, offsets them by the waist socket, and counts how many land on opaque torso pixels. Under 32 and the pose fails. Every view, every clip, every frame, every facing, every gait: 17,112 combinations.&lt;/p&gt;
&lt;p&gt;That catches the failure mode that makes sprites look like a collage, which is a regenerated part that no longer meets the part next to it.&lt;/p&gt;
&lt;h2&gt;So I tried to break it&lt;/h2&gt;
&lt;p&gt;Claiming you have tests is cheap. I reimplemented the audit standalone, outside the engine, then fed it deliberately corrupted atlases to find out what it actually catches.&lt;/p&gt;
&lt;p&gt;Baseline reproduces: 17,112 combinations, zero failures.&lt;/p&gt;
&lt;p&gt;Then the damage:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Character 4% smaller: &lt;strong&gt;FAIL, 352 bad combinations&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Character 8% smaller: &lt;strong&gt;FAIL, 6,728&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Alpha edge eroded 2px: &lt;strong&gt;FAIL, 3,036&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Torso shifted 6px down in the cell: &lt;strong&gt;PASS&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Scale drift gets caught at 4% and up, and slips through at 3%. Vertical translation is never caught. I pushed it to 14px and it still passed, because shifting the torso down actually &lt;em&gt;raises&lt;/em&gt; the minimum overlap from 59 to 249. The check gets happier while the sprite gets worse.&lt;/p&gt;
&lt;p&gt;That is a real hole, and it is the honest version of &amp;quot;we have tests.&amp;quot; The seam audit measures whether two parts still meet. It says nothing about whether the assembled pair sits correctly inside its cell, because the waist socket moves with the torso. Catching that needs a different check, anchored to the cell rather than to the neighbouring sprite.&lt;/p&gt;
&lt;p&gt;I would rather know what the test misses than assume it catches everything.&lt;/p&gt;
&lt;p&gt;The standalone auditor is &lt;code&gt;tools/seam_audit.py&lt;/code&gt; in the repo. Plain Python, PIL and numpy, no Godot required, so you can run the numbers above yourself.&lt;/p&gt;
&lt;h2&gt;What I&amp;#39;d tell you to steal from my repo&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Color-key over alpha. Deterministic beats convenient.&lt;/li&gt;
&lt;li&gt;Pin the grid numerically in the prompt and crop on those exact numbers.&lt;/li&gt;
&lt;li&gt;Name every animation phase. Never say &amp;quot;walk cycle&amp;quot; and hope.&lt;/li&gt;
&lt;li&gt;Baseline your seams and diff them. This is the whole safety net.&lt;/li&gt;
&lt;li&gt;Keep source atlases and prompts in the repo. Regeneration is a build step, so its inputs are source code.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The way this generalizes: if AI generation is nondeterministic, everything downstream has to be strict. The prompt is a spec that gets compiled by the importer, and CI runs the seam tests. Without that you don&amp;#39;t have a pipeline, at best it&amp;#39;s a slot machine you keep pulling until the art looks okay.&lt;/p&gt;
&lt;p&gt;&lt;img src=&quot;../../assets/hcr2/iron-belly.jpg&quot; alt=&quot;Iron Belly: securing a coolant pump while Orks contest the area&quot;&gt;&lt;/p&gt;
&lt;h2&gt;Try it&lt;/h2&gt;
&lt;p&gt;Mac and Android builds are on &lt;a href=&quot;https://github.com/nonatofabio/hive-city-rampage-ii&quot;&gt;GitHub Releases&lt;/a&gt;. They&amp;#39;re previews: macOS is ad-hoc signed and not notarized, Android uses a debug key, so both will warn you. Source builds with &lt;code&gt;make run&lt;/code&gt; if you have Godot 4.3.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Also published on &lt;a href=&quot;https://dev.to/nonatofabio_28/ai-generated-pixel-art-needs-a-build-system-not-better-prompts-280c&quot;&gt;DEV&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>The Certainty We Never Had</title>
    <link href="https://nonatofabio.github.io/blog/posts/certainty_we_never_had.html" rel="alternate"/>
    <id>https://nonatofabio.github.io/blog/posts/certainty_we_never_had.html</id>
    <published>2026-08-21T00:00:00Z</published>
    <updated>2026-08-21T00:00:00Z</updated>
    <summary>Software was never as predictable as we told ourselves. AI didn't break the promise. It exposed that the promise was always fiction.</summary>
    <category term="software-engineering"/>
    <category term="ai"/>
    <category term="sre"/>
    <category term="reliability"/>
    <category term="operations"/>
    <content type="html">&lt;div class=&quot;listen-box&quot; id=&quot;listen&quot;&gt;
  &lt;p class=&quot;listen-label&quot;&gt;🎧 Prefer to listen? 10:44, narrated locally on my homelab with &lt;a href=&quot;https://github.com/nonatofabio&quot;&gt;LVNA&lt;/a&gt;.&lt;/p&gt;
  &lt;audio id=&quot;post-audio&quot; controls preload=&quot;none&quot; style=&quot;width:100%;&quot;&gt;
    &lt;source src=&quot;../artifacts/the-certainty-we-never-had.mp3&quot; type=&quot;audio/mpeg&quot;&gt;
    Your browser doesn't support audio playback.
    &lt;a href=&quot;../artifacts/the-certainty-we-never-had.mp3&quot;&gt;Download the MP3&lt;/a&gt;.
  &lt;/audio&gt;
&lt;/div&gt;&lt;style&gt;
.listen-box{scroll-margin-top:2rem;border:1px solid rgba(128,128,128,.25);border-radius:.6rem;padding:1rem 1.1rem;margin:1.5rem 0;}
.listen-box .listen-label{margin:0 0 .6rem;font-size:.95rem;opacity:.8;}
&lt;/style&gt;&lt;script&gt;
(function(){
  var a=document.getElementById('post-audio');
  if(!a||typeof gtag!=='function')return;
  var fired={};
  function once(n){if(fired[n])return;fired[n]=true;gtag('event',n,{post:'certainty_we_never_had',surface:'post'});}
  a.addEventListener('play',function(){once('audio_play');});
  a.addEventListener('timeupdate',function(){
    if(!a.duration)return;
    var p=a.currentTime/a.duration;
    if(p&gt;=.25)once('audio_25');
    if(p&gt;=.5)once('audio_50');
    if(p&gt;=.75)once('audio_75');
  });
  a.addEventListener('ended',function(){once('audio_complete');});
  // Arriving via #listen (e.g. the audio link on LinkedIn) should land on the
  // player rather than the top of a 1,400-word essay.
  if(location.hash==='#listen'){
    a.preload='metadata';
    setTimeout(function(){a.scrollIntoView({block:'center'});},0);
  }
})();
&lt;/script&gt;&lt;p&gt;There&amp;#39;s a story our industry tells itself: traditional software is deterministic, therefore predictable, therefore safe. And AI, because it samples from a probability distribution, is none of those things.&lt;/p&gt;
&lt;p&gt;The story is comforting. It&amp;#39;s also wrong on both ends. It flatters conventional software with a certainty it never had, and it condemns AI against a baseline that exists only in textbooks.&lt;/p&gt;
&lt;p&gt;A note on where I&amp;#39;m arguing from. I&amp;#39;ve spent a long time on the operations side of large-scale cloud systems, and most of what convinced me of this sits in incident reviews I can&amp;#39;t quote. So I&amp;#39;m not going to pretend a citation is doing work that experience is doing. What follows is grounded in public literature wherever the literature will carry it, and flagged as my own judgment where it won&amp;#39;t. You should discount the second part accordingly, but I&amp;#39;d rather you discount it than have me dress it up.&lt;/p&gt;
&lt;h2&gt;The determinism fetish&lt;/h2&gt;
&lt;p&gt;Individual machine instructions are deterministic: same state, same input, same output. But we quietly upgrade that narrow mathematical fact into a much grander claim: that the deployed system will do what people intended.&lt;/p&gt;
&lt;p&gt;That inference doesn&amp;#39;t hold. Determinism guarantees consistency, and a program can be consistently wrong. Deterministic malware is still malware. A deterministic deadlock reliably deadlocks. Correctness is a relationship between what the code does and what people wanted, and determinism says nothing about the second half.&lt;/p&gt;
&lt;p&gt;This isn&amp;#39;t new. DeMillo, Lipton, and Perlis argued in 1979 that even formal proofs of programs are social artifacts. They establish confidence, not truth. NIST&amp;#39;s verification guidance still draws the same line: verification shows conformance to a specification. It cannot show the specification was right.&lt;/p&gt;
&lt;h2&gt;The hidden state is people&lt;/h2&gt;
&lt;p&gt;The deeper failure is operational, and it shows up somewhere most people don&amp;#39;t look.&lt;/p&gt;
&lt;p&gt;Ask why cloud operations is so expensive. The intuitive answer is hardware failure and the economics of scale, things nobody ever claimed were deterministic. I used to think that too. I don&amp;#39;t anymore, and the reason is the part I find hard to argue in public without showing you incident data I don&amp;#39;t own.&lt;/p&gt;
&lt;p&gt;Here&amp;#39;s what I can say. Hardware failure is the &lt;em&gt;solved&lt;/em&gt; part of operations. It has clean failure models and automated remediation; a disk dying is a scheduled inconvenience. The expensive, unautomatable part is sociotechnical: changes people make to live systems, and the reasoning behind those changes going missing.&lt;/p&gt;
&lt;p&gt;Peter Naur named the mechanism in 1985: programming is theory building. The program is not the theory. The theory (why that timeout is thirty seconds, what invariant that strange conditional protects, which dependency breaks if you touch this) lives in the heads of the people who built it. When they leave through reorgs, layoffs, or ordinary churn, the code stays perfectly deterministic while the system becomes &lt;em&gt;less predictable&lt;/em&gt;, because nobody can reconstruct the assumptions the code encodes.&lt;/p&gt;
&lt;p&gt;The team is now operating a system whose real specification exists nowhere.&lt;/p&gt;
&lt;p&gt;The only knowledge that survives this decay is what&amp;#39;s been compressed into hardened, slow-moving foundations like TCP, POSIX, and retry-with-backoff, the first principles that took decades to solidify. Everything above that layer is perishable. A large fraction of operations spend is the recurring cost of that perishability.&lt;/p&gt;
&lt;p&gt;So the hidden state that makes deployed software unpredictable isn&amp;#39;t just timing, faults, and adversaries. It&amp;#39;s epistemic. Retries, circuit breakers, canaries, chaos testing, on-call rotations, root-cause reviews. These aren&amp;#39;t decorations on a certain system. They&amp;#39;re the tribute an uncertain system pays to reality. And a good part of that uncertainty is an organization forgetting its own reasons.&lt;/p&gt;
&lt;h2&gt;The safe-baseline fallacy&lt;/h2&gt;
&lt;p&gt;Once you see that, the standard case against AI loses its footing. That case compares the &lt;em&gt;idealized abstraction&lt;/em&gt; of conventional software against the &lt;em&gt;observed production behavior&lt;/em&gt; of AI. The fair comparison is observed system versus observed system, under matched tasks, authority, and safeguards.&lt;/p&gt;
&lt;p&gt;Nancy Leveson has spent a career on the underlying point: safety is not a property of a component, human or mechanical. It&amp;#39;s emergent from the whole sociotechnical system: architecture, environment, authority, and recovery mechanisms together. A program doesn&amp;#39;t become safe because its source is deterministic. A human doesn&amp;#39;t become safe because they possess judgment.&lt;/p&gt;
&lt;p&gt;The right question is never &amp;quot;is this component certain?&amp;quot; It&amp;#39;s &amp;quot;what&amp;#39;s the probability of harm from this system, in this envelope, and is that lower or higher than what we run today?&amp;quot; The incumbent shouldn&amp;#39;t win merely by being incumbent.&lt;/p&gt;
&lt;p&gt;Richard Cook&amp;#39;s &lt;em&gt;How Complex Systems Fail&lt;/em&gt; explains why the incumbent feels safe anyway. After every incident, the team reconstructs a tidy causal chain, fixes it, and walks away feeling the event was predictable all along. It wasn&amp;#39;t. Post-hoc explainability is not ex ante predictability. Root-cause analysis is a &lt;em&gt;selected narrative&lt;/em&gt;, and its psychological side effect is renewed, unearned confidence in the deterministic story, right until the next failure arrives through a chain nobody drew.&lt;/p&gt;
&lt;p&gt;If you&amp;#39;ve sat through enough of these, you already know the feeling I&amp;#39;m describing.&lt;/p&gt;
&lt;h2&gt;Where the symmetry ends&lt;/h2&gt;
&lt;p&gt;Here&amp;#39;s where I part ways with the enthusiastic version of my own argument. Demolishing a fake baseline doesn&amp;#39;t make AI as safe as what it replaces. Symmetric burden of proof is not symmetric ease of proof, and three asymmetries survive the correction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Calibration.&lt;/strong&gt; We have centuries of actuarial data on how people fail, and institutions (liability, licensure, sanction) built to shape those failures. For frontier models we have evaluations that are weak proxies for deployment behavior, systems that shift under context and provider updates, and attack classes like prompt injection with no patch-and-converge story the way memory-safety bugs had. We can write the risk comparison as an equation; we can&amp;#39;t yet measure both sides with equal confidence.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Correlation.&lt;/strong&gt; Human error is buffered by cognitive diversity. A thousand engineers make a thousand different mistakes. A fleet of agents running one model can make the &lt;em&gt;same&lt;/em&gt; mistake, everywhere, simultaneously, from a single update. That&amp;#39;s monoculture risk at a scale with no good human analogue. Lamport&amp;#39;s fault-tolerance results are a warning here, not a comfort: they hold only under explicit assumptions about how many components fail independently.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Tempo.&lt;/strong&gt; AI acts faster than oversight loops built for human speeds. A bad decision that a review meeting would have caught next Tuesday can execute ten thousand times by Tuesday.&lt;/p&gt;
&lt;p&gt;Against these stand real advantages the old system never offered: complete execution traces, cheap replay, mechanically bounded permissions, adversarial testing at scale, and replacement without an HR process.&lt;/p&gt;
&lt;p&gt;Neither column wins by default. An AI-augmented system may be safer than the human workflow in one envelope and far more dangerous in another. Which is exactly why this gets settled empirically, per deployment, not by reference class.&lt;/p&gt;
&lt;h2&gt;The position&lt;/h2&gt;
&lt;p&gt;Software assurance has always been probabilistic at the system level. Determinism was a property of the abstraction. The certainty was manufactured by hindsight, sustained by institutional habit, and quietly financed by an operations budget whose real job was absorbing the gap between the model and the world, a gap widened every year by the churn that carries a system&amp;#39;s theory out the door.&lt;/p&gt;
&lt;p&gt;AI didn&amp;#39;t introduce uncertainty into software. It moved uncertainty somewhere we can no longer pretend not to see it.&lt;/p&gt;
&lt;p&gt;That relocation is uncomfortable, and it should be used well. The right response is neither to bar probabilistic components from serious work by comparing them to a fiction, nor to wave every agent through because &amp;quot;software was never certain anyway.&amp;quot;&lt;/p&gt;
&lt;p&gt;It&amp;#39;s to state, for every consequential system, whether human, deterministic, probabilistic, or mixed, the same disciplined claim: bounded probability of harm, explicit operating envelope, named assumptions, measured detection and recovery, and evidence the claim stays calibrated after deployment. Proofs, tests, monitors, and human escalation all become forms of evidence toward that claim, not talismans against needing one.&lt;/p&gt;
&lt;p&gt;The determinism fetish gave us permission to skip that discipline for fifty years. AI has revoked the permission.&lt;/p&gt;
&lt;p&gt;That may turn out to be its most valuable contribution to software engineering, delivered before anyone decides whether to trust it with anything else.&lt;/p&gt;
&lt;hr&gt;
&lt;h2&gt;Sources&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;R. DeMillo, R. Lipton, A. Perlis, &amp;quot;Social Processes and Proofs of Theorems and Programs&amp;quot; (CACM, 1979)&lt;/li&gt;
&lt;li&gt;P. Naur, &amp;quot;Programming as Theory Building&amp;quot; (1985)&lt;/li&gt;
&lt;li&gt;N. Leveson, &lt;em&gt;Engineering a Safer World&lt;/em&gt; (MIT Press, 2011)&lt;/li&gt;
&lt;li&gt;R. Cook, &amp;quot;How Complex Systems Fail&amp;quot; (1998)&lt;/li&gt;
&lt;li&gt;L. Lamport, R. Shostak, M. Pease, &amp;quot;The Byzantine Generals Problem&amp;quot; (1982)&lt;/li&gt;
&lt;li&gt;NIST guidance on software verification and validation&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;Views are my own and don&amp;#39;t represent my employer. Nothing here draws on non-public information.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>I Built My Own AI Agent (And Open-Sourced It)</title>
    <link href="https://nonatofabio.github.io/blog/posts/luna_agent.html" rel="alternate"/>
    <id>https://nonatofabio.github.io/blog/posts/luna_agent.html</id>
    <published>2026-03-04T00:00:00Z</published>
    <updated>2026-03-04T00:00:00Z</updated>
    <summary>Why I rejected every agent framework and built Luna — a ~2,300-line Python AI agent with SQLite hybrid-search memory, MCP tools, and Discord, fully local.</summary>
    <category term="ai"/>
    <category term="agents"/>
    <category term="open-source"/>
    <category term="python"/>
    <category term="homelab"/>
    <content type="html">&lt;p&gt;This started because I wanted a Discord bot that could remember things.&lt;/p&gt;
&lt;p&gt;Not a chatbot — I have plenty of those. I wanted an agent that could hold a conversation across days, search the web, run shell commands, and actually &lt;em&gt;learn&lt;/em&gt; who I am over time. The kind of thing where you message it on Tuesday about a project and on Friday it remembers the context without you re-explaining everything.&lt;/p&gt;
&lt;p&gt;So I went looking for a framework. That&amp;#39;s where the trouble started.&lt;/p&gt;
&lt;h2&gt;The Framework Problem&lt;/h2&gt;
&lt;p&gt;There are approximately ten thousand AI agent frameworks right now, and every week someone on Hacker News launches a new one. They all promise the same thing: &amp;quot;Build powerful AI agents in minutes!&amp;quot; I evaluated three seriously.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Framework&lt;/th&gt;
&lt;th&gt;Size&lt;/th&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;OpenClaw&lt;/td&gt;
&lt;td&gt;~400,000 lines&lt;/td&gt;
&lt;td&gt;42,000 exposed instances on Shodan. Impossible to audit.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ZeroClaw&lt;/td&gt;
&lt;td&gt;~2,000 lines&lt;/td&gt;
&lt;td&gt;9 days old. No community, uncertain future.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NanoClaw&lt;/td&gt;
&lt;td&gt;~500 lines&lt;/td&gt;
&lt;td&gt;Too thin. Missing memory, MCP, observability. Would rebuild most of it anyway.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The core needs — LLM chat, persistent memory, tool calling, Discord interface, structured logging — are individually simple and well-understood problems. No 400K-line framework needed. The risk of building from scratch was spending a few days writing Python. The risk of a framework was inheriting its complexity, its security surface, and its opinions about how agents should work.&lt;/p&gt;
&lt;p&gt;Easy tradeoff.&lt;/p&gt;
&lt;h2&gt;What Luna Actually Does&lt;/h2&gt;
&lt;p&gt;Luna is a custom AI agent that runs entirely on local hardware — two RTX 3090s (48GB total VRAM) running Qwen3-Coder-Next via llama-server, with no cloud APIs and no ongoing costs. The architecture is deliberately boring:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Discord (discord.py)
     |
     v
+-----------------------+
|    Luna Agent Core    |
|                       |
|  agent.py             |  agent loop: msg -&amp;gt; memory -&amp;gt; prompt -&amp;gt; LLM -&amp;gt; tools -&amp;gt; respond
|    +-- llm.py         |  single LLM client, configurable endpoint
|    +-- memory.py      |  SQLite + FTS5 + sqlite-vec hybrid search
|    +-- tools.py       |  native tools: bash, files, web fetch, web search
|    +-- tool_output.py |  smart output pipeline for large results
|    +-- mcp_manager.py |  MCP client for community tool servers
|    +-- observe.py     |  structured JSON logging
|                       |
+-----------------------+
         |
         v
   llama-server          any OpenAI-compatible endpoint
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;One process, one database file, one config file, eight runtime dependencies. The whole thing is ~2300 lines of Python including tests. Every line is auditable because there aren&amp;#39;t that many lines to audit.&lt;/p&gt;
&lt;h2&gt;The Design Choices That Matter&lt;/h2&gt;
&lt;p&gt;Building from scratch means you own every decision, which is both the privilege and the burden. Here are the ones that shaped Luna the most:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;SQLite for everything.&lt;/strong&gt; Messages, memories, full-text search, and vector search all live in a single file. FTS5 is built into SQLite, and sqlite-vec adds vector search without needing a separate vector database. The entire memory system backs up with &lt;code&gt;cp&lt;/code&gt;. I spent exactly zero hours configuring Postgres, managing Redis, or debugging connection pools. For a single-user agent running on a homelab, this is exactly the right level of infrastructure.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hybrid search with Reciprocal Rank Fusion.&lt;/strong&gt; Memory retrieval combines FTS5 keyword search with sqlite-vec semantic search. Keyword search catches the exact matches that embeddings miss — things like &amp;quot;error code E1234.&amp;quot; Vector search catches the semantic matches that keywords miss — &lt;em&gt;&amp;quot;the bug where the server crashes&amp;quot;&lt;/em&gt; finds a memory about a segfault even though neither word appears. RRF fuses the two result sets with one line of math per result, no trained model needed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One LLM endpoint, firewall-ready.&lt;/strong&gt; All traffic flows through a single &lt;code&gt;LLMClient&lt;/code&gt; pointing at one configurable URL. Today that&amp;#39;s &lt;code&gt;localhost:8001&lt;/code&gt;. When I&amp;#39;m ready to add an AI firewall — an input/output filtering proxy — I change the URL to &lt;code&gt;localhost:9000&lt;/code&gt; and put the proxy in front of the real LLM. Zero code changes required. I didn&amp;#39;t build the firewall, but I didn&amp;#39;t block the insertion point either.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Conversation compression.&lt;/strong&gt; Every 50 messages, the LLM summarizes the conversation and extracts facts with importance scores. Important facts go into long-term memory, and the summary keeps the conversation coherent across sessions. This gives effectively infinite conversation length — the agent always has context, even if the verbatim messages were summarized away days ago.&lt;/p&gt;
&lt;h2&gt;What I Didn&amp;#39;t Build (On Purpose)&lt;/h2&gt;
&lt;p&gt;Honestly, this is the list I&amp;#39;m most proud of:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;No AI firewall &lt;em&gt;(future — just don&amp;#39;t block the insertion point)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;No web dashboard &lt;em&gt;(the structured logs are ready for one when I want it)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;No multi-user auth &lt;em&gt;(single user)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;No cloud LLM fallback &lt;em&gt;(local only)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;No Docker &lt;em&gt;(systemd is simpler for a single-user Python process)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;No abstractions for hypothetical future requirements&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Every feature I &lt;em&gt;didn&amp;#39;t&lt;/em&gt; build is a feature I don&amp;#39;t have to maintain, secure, or debug at 2am. Sophisticated is the enemy of simple, complex is the enemy of valuable. The right amount of complexity is the minimum needed for the current problem.&lt;/p&gt;
&lt;h2&gt;How Memory Actually Works&lt;/h2&gt;
&lt;p&gt;This is the part I&amp;#39;m most technically proud of, because it&amp;#39;s the thing that makes Luna feel like more than a stateless chatbot. The memory system has three layers that work together:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Message history&lt;/strong&gt; is the raw conversation, stored per session. This is what gives the agent short-term context — the last 20 messages in the current thread.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Extracted memories&lt;/strong&gt; are facts the LLM identifies as worth remembering, each with an importance score from 1 to 10. A score of 10 (&lt;em&gt;&amp;quot;user&amp;#39;s name is Fabio&amp;quot;&lt;/em&gt;) always surfaces when relevant. A score of 2 (&lt;em&gt;&amp;quot;the weather was nice&amp;quot;&lt;/em&gt;) fades quickly. These persist across sessions — the agent builds a growing understanding of the world over time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Session summaries&lt;/strong&gt; are LLM-generated compressions of old message blocks. When the conversation gets long, older messages get summarized so the agent retains the gist without eating the entire context window.&lt;/p&gt;
&lt;p&gt;When the agent needs to recall something, search combines all three signals:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;final_score = rrf_score + (recency_weight * 2^(-age_days / 7)) + (importance / 10 * 0.1)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The recency decay has a 7-day half-life — a week-old memory scores half as much as a fresh one. All of these parameters live in &lt;code&gt;config.toml&lt;/code&gt;, so you can experiment with the tradeoffs without touching code.&lt;/p&gt;
&lt;p&gt;For embeddings, I went with nomic-embed-text-v1.5: 22M parameters, loads in seconds, runs entirely on CPU without touching GPU memory. It supports Matryoshka representations, which means I can use 384 dimensions now and scale to 768 later without re-embedding everything.&lt;/p&gt;
&lt;h2&gt;Tools: Native and MCP&lt;/h2&gt;
&lt;p&gt;Luna ships with six built-in tools — bash (with safety guardrails), file read/write, directory listing, web fetch, and web search. The bash tool checks commands against blocked patterns before execution, enforces a 30-second timeout, and caps output at 50KB. No &lt;code&gt;rm -rf /&lt;/code&gt;, no &lt;code&gt;mkfs&lt;/code&gt;, no fork bombs.&lt;/p&gt;
&lt;p&gt;For everything beyond the builtins, there&amp;#39;s MCP. The Model Context Protocol lets you connect community tool servers by editing a JSON config:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-json&quot;&gt;{
  &amp;quot;servers&amp;quot;: {
    &amp;quot;browser&amp;quot;: {
      &amp;quot;command&amp;quot;: &amp;quot;npx&amp;quot;,
      &amp;quot;args&amp;quot;: [&amp;quot;-y&amp;quot;, &amp;quot;@playwright/mcp&amp;quot;],
      &amp;quot;transport&amp;quot;: &amp;quot;stdio&amp;quot;
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Need browser automation? Add the Playwright MCP server. Need filesystem tools? Add that server. Each one runs as a separate process with natural isolation, and tool names get namespaced automatically (&lt;code&gt;browser__navigate&lt;/code&gt;, &lt;code&gt;filesystem__read_file&lt;/code&gt;) so nothing collides. Adding a new capability to the agent is editing JSON, not writing code.&lt;/p&gt;
&lt;h2&gt;The Numbers&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;~2300 lines&lt;/strong&gt; of Python (agent + tests)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;106 tests&lt;/strong&gt;, all passing — no GPU or running LLM required to run them&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;8 runtime dependencies&lt;/strong&gt; — discord.py, openai, mcp, sentence-transformers, einops, sqlite-vec, html2text, duckduckgo-search&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;~187 tokens/sec&lt;/strong&gt; prompt processing, &lt;strong&gt;~81 tokens/sec&lt;/strong&gt; generation on 2x RTX 3090&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Why Open Source It&lt;/h2&gt;
&lt;p&gt;I built Luna for my homelab, and it works well for what I need. But the decisions behind it — custom build over framework, SQLite over Postgres, local LLM over cloud API, deliberate simplicity over feature checklists — those aren&amp;#39;t unique to my setup. Anyone with a GPU and a desire to actually &lt;em&gt;understand&lt;/em&gt; their AI agent stack could use this as a starting point, or at least steal the ideas they like.&lt;/p&gt;
&lt;p&gt;The code is MIT licensed. The &lt;a href=&quot;https://github.com/nonatofabio/luna-agent/blob/main/DESIGN.md&quot;&gt;DESIGN.md&lt;/a&gt; explains the reasoning behind every major architectural decision. The tests mock the LLM client so you can run them on a laptop. The config is a single TOML file with sensible defaults.&lt;/p&gt;
&lt;p&gt;If you&amp;#39;re tired of agent frameworks that are bigger than the applications they power, maybe start with something you can actually read.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://github.com/nonatofabio/luna-agent&quot;&gt;GitHub: nonatofabio/luna-agent&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;The repo has a few &lt;a href=&quot;https://github.com/nonatofabio/luna-agent/issues?q=is%3Aopen+label%3A%22good+first+issue%22&quot;&gt;good first issues&lt;/a&gt; if you want to contribute. And if you build something interesting with it, I&amp;#39;d genuinely love to hear about it.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Thanks for reading.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Keep it Awesome!&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>I Co-Authored a Research Paper With an AI Agent</title>
    <link href="https://nonatofabio.github.io/blog/posts/ai_coauthor.html" rel="alternate"/>
    <id>https://nonatofabio.github.io/blog/posts/ai_coauthor.html</id>
    <published>2026-02-04T00:00:00Z</published>
    <updated>2026-02-04T00:00:00Z</updated>
    <summary>How I ran an ML research project with an AI agent as a real collaborator — GaLore continual-learning experiments, a rank scaling law, and a co-authored paper.</summary>
    <category term="ai"/>
    <category term="ml"/>
    <category term="research"/>
    <category term="agents"/>
    <category term="collaboration"/>
    <content type="html">&lt;p&gt;Look, I&amp;#39;m going to be honest with you. This whole thing started as an experiment within an experiment.&lt;/p&gt;
&lt;p&gt;The &lt;em&gt;official&lt;/em&gt; research question was: &amp;quot;Can a language model learn continuously from human conversations without forgetting everything it already knows?&amp;quot; But the &lt;em&gt;real&lt;/em&gt; experiment? Whether I could run an entire ML research project - from hypothesis to paper - with an AI agent as my co-pilot. Not as a fancy autocomplete. As an actual collaborator.&lt;/p&gt;
&lt;p&gt;Spoiler: It worked. &lt;a href=&quot;../artifacts/paper.pdf&quot;&gt;We wrote a paper&lt;/a&gt;. And the process was, aham, &lt;em&gt;weird&lt;/em&gt;.&lt;/p&gt;
&lt;h2&gt;The Setup: Human + Agent = ???&lt;/h2&gt;
&lt;p&gt;Here&amp;#39;s how it worked. I&amp;#39;m a scientist with 15+ years in ML. I had some research intuition, the domain knowledge, and access to 8x A100 GPUs on a remote server. What I didn&amp;#39;t have was infinite time to write boilerplate code, babysit training runs, and manually parse log files at 2am.&lt;/p&gt;
&lt;p&gt;Enter the agent.&lt;/p&gt;
&lt;p&gt;My AI collaborator could execute bash commands, write and modify code, SSH into the training server, run experiments, parse results, and—critically: &lt;em&gt;remember context across our entire conversation&lt;/em&gt;. It wasn&amp;#39;t just answering questions. It was maintaining state. Tracking what we&amp;#39;d tried. Suggesting next steps.&lt;/p&gt;
&lt;p&gt;The workflow looked like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Me: &amp;quot;Let&amp;#39;s try reducing the GaLore rank to 64 and see if it helps with forgetting&amp;quot;
Agent: *writes config file*
Agent: *SSHs to server*
Agent: *launches training run*
Agent: *monitors logs*
Agent: &amp;quot;Training complete. MMLU dropped 3.0% vs 4.8% before. Want me to run holdout eval?&amp;quot;
Me: &amp;quot;Yes&amp;quot;
Agent: *runs eval, parses results, updates experiment notes*
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I was the advisor. The agent was the grad student who never sleeps and never complains about running &amp;quot;just one more ablation.&amp;quot;&lt;/p&gt;
&lt;h2&gt;Day 1: The Naive Approach (We Both Got It Wrong)&lt;/h2&gt;
&lt;p&gt;Our first attempt was embarrassingly simple. Train a 0.5B model on conversation data. Check if it learned. Check if it forgot. The agent set everything up: data pipelines, training loop, GaLore optimizer. I reviewed the code, made some suggestions, and we kicked off a 500-step run.&lt;/p&gt;
&lt;p&gt;The training loss looked beautiful. Smooth curves. Decreasing numbers. Then we ran the benchmarks.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;MMLU: 48.26% → 43.45% (-4.8%)
Holdout perplexity: +14% to +22% WORSE
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;strong&gt;The model got dumber.&lt;/strong&gt; Classic overfitting. We&amp;#39;d both missed it.&lt;/p&gt;
&lt;p&gt;Here&amp;#39;s the thing: the agent didn&amp;#39;t try to hide the failure or spin it. It just... reported the results and asked what we should try next. No ego. No defensiveness. Just &amp;quot;well, that didn&amp;#39;t work. Here are some hypotheses.&amp;quot; That&amp;#39;s when I realized this collaboration might actually work.&lt;/p&gt;
&lt;h2&gt;The Design of Experiments: Where the Agent Earned Its Keep&lt;/h2&gt;
&lt;p&gt;I decided we needed a proper factorial experiment. Three hyperparameters, eight combinations, run in parallel. The kind of thing that&amp;#39;s conceptually simple but logistically annoying.&lt;/p&gt;
&lt;p&gt;Me: &amp;quot;Let&amp;#39;s do a 2³ DOE. Factors are rank reduction, LR scheduler, and weight decay.&amp;quot; The agent:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Generated all 8 config files&lt;/li&gt;
&lt;li&gt;Wrote a bash script to launch them in parallel across GPUs&lt;/li&gt;
&lt;li&gt;Wrote an eval script to benchmark all 8 models&lt;/li&gt;
&lt;li&gt;Created a results table in our experiment notes&lt;/li&gt;
&lt;li&gt;Ran the whole thing while I went to get coffee&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;[Insert Code Snippet: The DOE launch script the agent wrote in about 30 seconds - please, Claude Intern 1.0, fix this!]&lt;/p&gt;
&lt;p&gt;When I came back, I had this:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Config&lt;/th&gt;
&lt;th&gt;Rank&lt;/th&gt;
&lt;th&gt;Scheduler&lt;/th&gt;
&lt;th&gt;Weight Decay&lt;/th&gt;
&lt;th&gt;MMLU&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;3.0&lt;/td&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;0.1&lt;/td&gt;
&lt;td&gt;44.15%&lt;/td&gt;
&lt;td&gt;Baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3.1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;64&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;❌&lt;/td&gt;
&lt;td&gt;0.1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;45.26%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;WINNER&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3.6&lt;/td&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;td&gt;✅&lt;/td&gt;
&lt;td&gt;0.2&lt;/td&gt;
&lt;td&gt;43.90%&lt;/td&gt;
&lt;td&gt;Worst&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;td&gt;...&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;The agent had already identified the pattern: &lt;strong&gt;rank was the dominant factor&lt;/strong&gt;. Scheduler hurt. Weight decay did nothing. &lt;em&gt;I didn&amp;#39;t have to parse a single log file.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;The Scaling Crisis: When We Hit a Wall&lt;/h2&gt;
&lt;p&gt;Feeling confident, we tried the winning config on bigger models. TinyLlama (1.1B) worked perfectly. Then Gemma-2 (2.6B) broke everything.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;MMLU drop: -5.67%
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The agent flagged it immediately: &amp;quot;This exceeds the 5% threshold. The rank-64 setting may be too permissive for larger models.&amp;quot; This is where the human-agent dynamic got interesting. The agent had the data. I had the intuition. Together, we hypothesized that bigger models need &lt;em&gt;lower&lt;/em&gt; rank, meaning more constraint, not less.&lt;/p&gt;
&lt;p&gt;Me: &amp;quot;Try rank=32 on Gemma&amp;quot;
Agent: &lt;em&gt;runs experiment&lt;/em&gt;
Agent: &amp;quot;MMLU now -3.85%. Within threshold.&amp;quot;&lt;/p&gt;
&lt;p&gt;We&amp;#39;d discovered a &amp;quot;scaling law&amp;quot;: &lt;code&gt;rank ∝ 1/√params&lt;/code&gt;. The agent helped me formalize it into a table:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model Size&lt;/th&gt;
&lt;th&gt;Optimal Rank&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;&lt;tr&gt;
&lt;td&gt;0.5B-1.1B&lt;/td&gt;
&lt;td&gt;64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2B-3B&lt;/td&gt;
&lt;td&gt;32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8B&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;&lt;/table&gt;
&lt;p&gt;Neither of us would have gotten there alone. I needed the agent to run the experiments fast enough to iterate. The agent needed me to recognize the pattern and propose the hypothesis.&lt;/p&gt;
&lt;h2&gt;The 8B Moment: When It Actually Worked&lt;/h2&gt;
&lt;p&gt;Time for the real test. Qwen3-8B. 8.2 billion parameters. Rank=8. 1000 training steps. The agent ran it overnight. I woke up to this message:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Training complete. 5:40 duration, 775 tokens/sec.
MMLU: 74.93% → 75.07% (+0.14%)
Holdout PPL: 5.34 → 4.87 (-8.8%)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Wait. &lt;strong&gt;The model got smarter?&lt;/strong&gt; I didn&amp;#39;t believe it. I asked the agent to run a statistical significance test.&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Paired t-test: t=7.12, p&amp;lt;0.0001
Significant: True
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It was real. The model learned from conversations &lt;em&gt;and&lt;/em&gt; improved on benchmarks. Not just &amp;quot;acceptable forgetting&amp;quot;, it got net positive knowledge transfer. The agent&amp;#39;s response: &amp;quot;This proves the core hypothesis. Want me to update the experiment notes and commit?&amp;quot;&lt;/p&gt;
&lt;p&gt;Yes. Yes I did.&lt;/p&gt;
&lt;h2&gt;The LoRA Showdown: A Plot Twist&lt;/h2&gt;
&lt;p&gt;I had a nagging question: how does this compare to LoRA, the thing everyone actually uses? The agent ran the comparison. Same model, same data, same steps, same rank.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Efficiency:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LoRA: 22GB VRAM, 982 tokens/sec ✅&lt;/li&gt;
&lt;li&gt;GaLore: 39GB VRAM, 775 tokens/sec&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Learning quality:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;LoRA holdout PPL: +530% (catastrophic failure)&lt;/li&gt;
&lt;li&gt;GaLore holdout PPL: -8.8% (genuine learning)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;LoRA was faster and lighter. It also &lt;em&gt;completely failed to learn anything generalizable&lt;/em&gt;. The model memorized training data and forgot how to generalize. The agent&amp;#39;s analysis was spot-on: &amp;quot;LoRA freezes base weights and only trains adapters. It can&amp;#39;t integrate new knowledge—only patch outputs. GaLore projects gradients but updates all weights, enabling genuine learning.&amp;quot;&lt;/p&gt;
&lt;p&gt;That insight made it into the paper almost verbatim.&lt;/p&gt;
&lt;h2&gt;Writing the Paper: The Final Boss&lt;/h2&gt;
&lt;p&gt;After six phases of experiments, we had results. Now we needed a paper. This is where I expected the collaboration to break down. Writing is &lt;em&gt;hard&lt;/em&gt;. It requires judgment, narrative, argumentation. Surely an AI can&amp;#39;t...&lt;/p&gt;
&lt;p&gt;The agent drafted the abstract in one shot. It was... good? Like, actually good. It captured the key findings, the methodology, the implications. I edited maybe 20%.&lt;/p&gt;
&lt;p&gt;We went section by section:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;I&amp;#39;d outline what needed to be said&lt;/li&gt;
&lt;li&gt;The agent would draft it&lt;/li&gt;
&lt;li&gt;I&amp;#39;d revise and push back&lt;/li&gt;
&lt;li&gt;The agent would incorporate feedback&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The related work section was particularly impressive. The agent pulled relevant citations, summarized them, and positioned our work in the literature. I added a few papers it missed, but the structure was solid.&lt;/p&gt;
&lt;p&gt;If you look at my &lt;code&gt;git log&lt;/code&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;eda30d8 exp006: Add statistical significance test - p&amp;lt;0.0001, t=7.12
0151460 exp006 Phase 3: GaLore vs LoRA comparison - GaLore wins
53a2c88 exp006 Phase 2: Add learning measurement
9b1e689 exp006 Phase 2: Extended training shows +0.14% MMLU improvement
f385da7 Update paper with exp006 8B validation
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Every commit was a collaboration. Every result was verified. Every claim was backed by an experiment we&amp;#39;d run together.&lt;/p&gt;
&lt;h2&gt;What I Learned About Human-Agent Collaboration&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;1. The agent is a force multiplier, not a replacement.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I couldn&amp;#39;t have run this many experiments this fast alone. But the agent couldn&amp;#39;t have designed the experiments or recognized the patterns without me. We were genuinely complementary.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Context is everything.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The agent remembered our entire conversation, every failed experiment, every hypothesis, every decision. It could say &amp;quot;remember when we tried X and it didn&amp;#39;t work because Y?&amp;quot; That continuity was invaluable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. The agent has no ego.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;When experiments failed, the agent just... moved on. No defensiveness. No excuses. Just &amp;quot;that didn&amp;#39;t work, here&amp;#39;s what we could try next.&amp;quot; It&amp;#39;s weirdly refreshing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. Trust but verify.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I checked every result. Every claim. Every number. The agent made mistakes, small ones, big ones. But the verification loop kept us honest.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;5. The meta-irony is not lost on me.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;We built a system for AI to learn from human conversations. We did it &lt;em&gt;through&lt;/em&gt; human-AI conversation. The research method mirrored the research question.&lt;/p&gt;
&lt;h2&gt;The Takeaway&lt;/h2&gt;
&lt;p&gt;The paper&amp;#39;s conclusion is about GaLore and continuous learning. But my conclusion is different.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;We&amp;#39;re entering an era where research itself can be collaborative between humans and AI agents.&lt;/strong&gt; Not AI replacing researchers. Not humans doing everything manually. Something in between - a partnership where each side contributes what they&amp;#39;re good at.&lt;/p&gt;
&lt;p&gt;I brought intuition, judgment, and experience. The agent brought speed, memory, and tireless execution. Together, we wrote a paper that neither of us could have written alone.&lt;/p&gt;
&lt;p&gt;The answer to &amp;quot;Can AI learn continuously from human conversations?&amp;quot; turned out to be yes.&lt;/p&gt;
&lt;p&gt;But the more interesting answer? &amp;quot;Can humans and AI agents do research together?&amp;quot;&lt;/p&gt;
&lt;p&gt;Also yes.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href=&quot;../artifacts/paper.pdf&quot;&gt;The paper is available here&lt;/a&gt;. The code we wrote is in: &lt;a href=&quot;https://github.com/nonatofabio/continuous-learning&quot;&gt;Continuous Learning&lt;/a&gt;. And I&amp;#39;m going to go touch grass now, something my co-author will never need to do.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Thanks for reading.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Keep it Awesome!&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
  <entry>
    <title>Welcome to My Blog</title>
    <link href="https://nonatofabio.github.io/blog/posts/welcome.html" rel="alternate"/>
    <id>https://nonatofabio.github.io/blog/posts/welcome.html</id>
    <published>2026-02-04T00:00:00Z</published>
    <updated>2026-02-04T00:00:00Z</updated>
    <summary>Why I'm starting a blog: practical lessons from AWS-scale ML infrastructure, developer tools like MCP servers, cybersecurity AI, and open-source deep dives.</summary>
    <category term="meta"/>
    <category term="introduction"/>
    <content type="html">&lt;p&gt;I&amp;#39;ve been meaning to come back to writing for a while now. After years of building tools, shipping features, and breaking AI/ML infrastructure, I realized I learned a lot of stuff that might be useful to share.&lt;/p&gt;
&lt;h2&gt;What to Expect&lt;/h2&gt;
&lt;p&gt;This blog will cover topics around my interests:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AI/ML Infrastructure&lt;/strong&gt; - Practical lessons from building and scaling ML systems at AWS&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Developer Tools&lt;/strong&gt; - Building useful things like MCP servers and local AI tooling&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cybersecurity&lt;/strong&gt; - Threat intelligence and security applications of machine learning&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open Source&lt;/strong&gt; - Deep dives into projects I&amp;#39;m working on&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Why Write?&lt;/h2&gt;
&lt;p&gt;Writing forces clarity. When you have to explain something, you find gaps in your own understanding. I&amp;#39;m hoping this blog helps me think more clearly, while maybe helping others along the way.&lt;/p&gt;
&lt;h2&gt;The Tech Stack&lt;/h2&gt;
&lt;p&gt;This blog is intentionally simple: Markdown files rendered client-side with marked.js, hosted on GitHub Pages. No build step, no database, no complexity. Just text files and a browser.&lt;/p&gt;
&lt;p&gt;Sometimes the best tool is the simplest one that gets the job done.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Stay tuned for more posts. You can find me on &lt;a href=&quot;https://github.com/nonatofabio&quot;&gt;GitHub&lt;/a&gt; or &lt;a href=&quot;https://www.linkedin.com/in/fabiononato/&quot;&gt;LinkedIn&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
</content>
  </entry>
</feed>
