25 rounds. 2 wins. One commit that moved performance from 57 to 74.

I run ratradar.com — a health inspection radar for US cities. It loads a Next.js frontend with Google Maps, Ahrefs analytics, and GTM. The Lighthouse score was sitting at 57 on performance. That's not catastrophic but it's not great either. So I decided to let an AI loop fix it for me and see what happened.

The idea

Auto-Lighthouse is simple: drop in a URL, the loop runs a Lighthouse audit, reads the results, generates a code change to improve the score, applies it, runs Lighthouse again, and keeps the change if the score went up. Repeat.

No human in the loop. No cherry-picking. The AI decides what to change, applies it, and lives with the result.

I gave it 25 rounds to work with.

What happened

Most rounds: nothing. The AI tried things that sounded reasonable in theory — lazy loading this, deferring that, tweaking cache headers — and the score moved by exactly 0.

Two rounds actually won.

The one that mattered: moving Google Tag Manager and Ahrefs to lazyOnload.

Both scripts were loading as afterInteractive — which means they blocked the main thread during the interactive phase. Switching them to lazyOnload drops them until the page is fully idle. That single commit moved performance from 57 to 74. Composite score went from 77 to 86.

What the AI got wrong

A lot.

It tried to optimize images that were already optimized. It suggested splitting bundles that were already being split by Next.js. It proposed cache header changes that didn't apply to our deployment setup. It was confidently wrong in exactly the way AI systems are confidently wrong — plausible-sounding changes that don't account for the actual runtime context.

The loop caught most of this automatically: apply the change, run Lighthouse, score didn't improve, revert. The experiment was limited to 25 rounds.

What's left

Two things are still dragging the score:

  • LCP 8.9s — the restaurant page renders a Google Maps WebGL component that dominates LCP. The fix isn't a script tag change, it's architectural — either lazy-load the map behind an interaction or rethink the above-fold layout.
  • CLS 0.329 — layout shifts happening during map load. Related to the same component.

These require targeted fixes, not a generic optimization loop. The AI loop hit its limit where the remaining gains require understanding the specific component, not just the script loading strategy.

The tool

The Auto-Lighthouse project page describes the experiment and its results.

The interesting thing is how well it works for low-hanging fruit. GTM and analytics script placement is something every site gets wrong and no one fixes because it requires knowing which Next.js loading strategy to use. The AI knew. It just needed a loop to find the right moment to apply it.

That's most of what automated optimization is — knowing the right answer, finding the right moment to apply it.