Chasing 100 Lighthouse scores on my Astro blog
One of the reasons I recently migrated to Astro was better architecture for performance improvements. After all, a blog’s performance doesn’t have to suck, right? It’s just a blog! Static files and a bunch of interactivity with JS I like to add.
The natural way to measure this and improve is to use Lighthouse. But be aware that measuring just the homepage might not be enough.
How should we do it in the AI era?
Build deterministic tools
To avoid costly AI runs fix after fix, we want AI to build a small harness that runs Lighthouse on every page, collects the results, and saves them.
Let’s do that first:
Build me a reusable Lighthouse harness for this Astro blog. I want a script Ican re-run after every change, not a one-off report.It should:- Audit the BUILT site (astro build output served by astro preview), never thedev server. Same setup on every run so results are comparable over time.- Discover every route automatically from the build output (homepage, posts,tools, 404 excluded) - no hardcoded list.- Run Lighthouse against each route and save the raw JSON report per route,plus one summary file with the key metrics per route: score, FCP, LCP, CLS,TBT, Speed Index, total byte weight. Machine-readable JSON, so you can parseit later.- Support a desktop preset by default and an optional mobile run. Focus on theperformance category.- Expose single commands: `yarn perf` for the homepage and `yarn perf:all` forevery route.- Keep the generated reports in git so I can diff them later.- Print a readable summary table to stdout when it finishes.Then run it once so I can see the baseline.
The plan and the loop
Once we have our harness, and the first results, we can start working on a plan and tell AI how to loop through it. We definitely want to have a human in the loop, so every action should be consulted with the operator. After all, it’s us deciding if something is worth pursuing or not.
I know that this one has a potential to loop autonomously, because we have a clear goal of getting a 100 Lighthouse score on every page, but believe me, understanding the tradeoffs and deciding on them is the engineering part we need.
Now we have a harness and a baseline. Let's turn that into a plan and workthrough it, one step at a time.First, read the raw Lighthouse JSON reports (not just the summary scores). Pullout the failing audits and opportunities - the diagnostics - and give me:- a short list of the real issues, ranked by impact,- grouped into what's shared across pages vs what's specific to a page or two,- for each, what you'd change and the tradeoff involved (bundle size, visualchange, first-load font, etc.).Don't fix anything yet - I want to agree on the plan first.Then we work through it iteratively, and for EVERY change:1. Propose the fix and the tradeoff, and wait for my go-ahead. Don't batchunrelated changes.2. Implement only that fix.3. Re-run the harness (`yarn perf:all`).4. Show me the before/after delta per route, and call out anything that gotworse or didn't move.5. Commit it on its own, then move to the next item.Rules:- Optimize for real issues, not the score. A 100 isn't the goal if it costsreadability or the site's character - flag those tradeoffs to me.- Prefer fixing shared root causes once (e.g. fonts, image sizing) over patchingeach page.- If a fix doesn't measurably help, say so and we drop it.- Pause for my input at every decision point.
On this website, if you let AI do extreme ownership on delivering 100 score, it would, for example, remove the extended variant of the font from the homepage. Unless it would notice “ż” in my last name[1]Tomek Nieżurawski to understand I need the extended variant to support this Polish character. I call it unlikely.
The results
I’ll get to what we’ve done later, but first… the results! Here’s the before/after across all pages:
| Metric | before | after | delta |
|---|---|---|---|
| Desktop avg score | 97.6 | 100.0 (all pages) | +2.4 |
| Mobile avg score | 94.0 | 99.3 | +5.3 |
| Desktop bytes | 11,195 KiB | 6,482 KiB | –4.6 MB |
| Mobile bytes | 12,619 KiB | 6,422 KiB | –6.1 MB |
| Routes with CLS > 0.1 | 5 | 0 | cleared |
The averages hide where it hurt the most. The pages that were shipping oversized images and a pile of JavaScript saw the biggest jumps:
| Route | Score | LCP | Bytes |
|---|---|---|---|
/sideprojects | 75 → 96 | 10.6s → 2.7s | 2,065 → 498 KiB |
/posts/effective-titles-and-more | 78 → 99 | 5.8s → 2.1s | 1,113 → 358 KiB |
/posts/benchmarking-graphql-solutions-in-the-js-ts-landscape | 84 → 98 | 4.3s → 2.3s | 1,412 → 393 KiB |
Now we are talking! Nice difference. But also not going too crazy to have 100 on every page.
What actually moved the needle
A lot of Lighthouse advice is noise, but a few fixes paid off big:
- Self-hosted fonts + preload
- Images, images, images 😅 - Added explicit
width/heighteverywhere to stop the layout jumping, putfetchpriority="high"on hero images, and served post images as responsive WebP through Astro’s<Picture> - Deleted dead JavaScript - An old Universal Analytics tag (
UA-..., sunset in 2023) was still loading its runtime on every page. - Deferred island hydration - Most React demos didn’t need to hydrate immediately, so they moved from
client:loadtoclient:visible. The react-dom renderer now only loads when a demo scrolls into view. - Trimmed heavy dependencies - Swapped a monolithic color library for
colord(42 KB → 11.4 KB) and replacedpako+base64-jswith nativeDecompressionStream(45 KB → 1 KB).
I guess for anyone in the Lighthouse performance business, those are not surprising things to see! But the more important thing is that all those fixes were applied semi-automatically.
Takeaways
Build a performance harness, plan, loop and give your engineering input.
It’s that simple.
- [1]Tomek Nieżurawski
If you'd like to see more content like this, please follow me on Twitter: tomekdev_. I usually write about interfaces in the real life. So where the design meets implementation.
Comments and discussion welcomed in this thread.
