Skip to content
← Back to blog
14 min readGEO

How We Made GazeRank Scans 150x Faster by Caching AI Results (Instead of Calling Gemini Every Time)

Every AI product starts the same way. You wire up the model, send it a prompt, get a response, ship it. The first version of GazeRank did exactly that: user submits a URL, we crawl the site, send the signals to Gemini, get back a report. It worked. It was slow — typically 25 to 30 seconds per scan — but it worked.

Then we did the math. And we should have done it on day one.

The problem with “just call the model”

A monitor in GazeRank re-scans a URL on a schedule. Take a simple example: 20 sites, checked hourly. That's 480 scans a day, or about 14,400 a month. At ~30 seconds of Gemini time each, you're paying for roughly 120 hours of model inference monthly. On the free tier, you're also hitting rate limits constantly.

The worst part: most of those scans return the same answer.

If a site owner updates their pricing page once a month, and you scan every hour, you're calling Gemini hundreds of times to ask the same question about the same content. The content hash hasn't changed. The signals haven't changed. The model's answer — if you're asking the same question about the same input — shouldn't change either.

We were paying for thousands of identical answers a month.

The insight that fixed it

The signal we send to Gemini is deterministic given the input. If the page content is byte-identical to the last time we scanned, and our prompt hasn't changed, the output is going to be functionally the same — same score, same issues, same summary.

So we stopped re-asking.

We now compute a hash of the visible text of every page we crawl:

def _content_hash(visible_text: str) -> str:
    normalized = " ".join(visible_text.split())
    return hashlib.sha256(normalized.encode("utf-8")).hexdigest()[:16]

Two things matter here:

  • Whitespace normalization. We collapse all runs of whitespace to a single space before hashing. Otherwise, a CMS adding a newline in a template would produce a different hash and burn a Gemini call for nothing.
  • Truncation to 16 hex chars. That's 64 bits of entropy — more than enough for collision resistance in this context — and it keeps the cache keys short.

How the cache key is built

The hash of a single page isn't enough. We need a key that represents the entire scan for a specific site. So we:

  1. Normalize the domain (host only, lowercased, www. stripped)
  2. Hash every page we crawled
  3. Sort those hashes (order-independent — same pages in different crawl order still match)
  4. Append a prompt version tag
  5. Hash the whole thing
def _make_cache_key(url: str, content_hashes: list[str]) -> str:
    domain = normalize_domain(url)  # e.g. example.com, not www.example.com
    normalized = sorted(h for h in content_hashes if h)
    joined = f"{domain}::{('|'.join(normalized))}::{PROMPT_VERSION}"
    return hashlib.sha256(joined.encode("utf-8")).hexdigest()

Three design decisions worth explaining

Domain in the key

www.example.com and example.com share a key when content matches. Different domains never share an AI result, even if content hashes collide. Without this, a cache hit could return another site's summary under your URL — a correctness bug, not just a performance miss.

Sorting

If the crawler finds pages A, B, C on Monday and C, A, B on Tuesday (threading order), both should hit the same entry. Sorting removes false misses.

PROMPT_VERSION

When we materially change the system prompt — new instructions, scoring rubric, or output shape — we bump this from v1 to v2. Every entry under the old version becomes unreachable. The next scan gets a fresh answer under the new prompt. Without that tag, we'd ship a prompt improvement and users could see stale results for hours while old entries lived on. It's the easiest bug in the world to ship.

Real-world cases where hashes collide (or nearly do)

The scary failures aren't cryptographic collisions — they're two different businesses with the same empty React shell.

Content hashing is only as good as what you hash. In production we kept hitting situations where different sites produced the same — or dangerously similar — inputs:

1. Empty or near-empty HTML (SPA shells)

Many marketing sites ship a nearly blank document and render in JavaScript:

<html><body><div id="root"></div></body></html>

Visible text is empty or a few shared strings (“Loading...”, cookie banner copy). Two unrelated products on the same framework can hash identically. If the AI cache key were only content hashes, site B could inherit site A's summary and score narrative. Domain in the key stops that.

2. Shared boilerplate and cookie banners

A lot of “content” is the same across customers of one CMS or consent vendor: identical footer legal blocks, “We use cookies” modals, newsletter CTAs. After whitespace normalization, the distinct part of the page can be a thin slice. Hashing still works for change detection on the same site, but it is a weak fingerprint across sites. Again: scope by domain.

3. Placeholder / coming-soon pages

“Coming soon”, “Under construction”, default Nginx/Apache pages, and parked domains often share the same few dozen characters. Collisions here are not theoretical — they're common on the open web.

4. Soft collisions from truncation

We use 16 hex characters (64 bits) of the SHA-256. Accidental collision probability for a modest cache is vanishingly small for random inputs. The risk we care about is not birthday-attack crypto theory — it's structured sameness: empty bodies, identical templates, mirrored staging copies. Domain isolation is the practical fix; longer digests don't help if the normalized text is the same.

5. Same content, different brand in the chrome

Title tags and nav labels can differ while the main visible copy is cloned (franchise sites, white-label SaaS, multi-tenant landing pages). Pure text hashes may match; the URL host is still the identity the report must use. That's why we both key by domain and force the summary to name the scanned host.

6. What we deliberately do not treat as a change

Whitespace-only edits, minor template newlines, and crawl-order differences are normalized away so we don't false-miss. A real copy edit, a new H1, or a removed section changes the normalized string → new hash → cache miss → fresh Gemini call. That's the intended miss.

The rule we settled on:
Same domain + same content hashes + same prompt version → reuse the AI result.
Anything else → call the model.

Hashing answers “did this site's content change?”
The domain answers “whose report is this?”
Mixing those two questions is how you get a fast cache that is wrong. Keeping them separate is how you get a fast cache you can ship.

What happens on a cache hit

In run_scan, before we touch the model:

content_hashes = [
    p.get("content_hash")
    for p in signals.get("pages", [])
    if p.get("content_hash")
]
cached_ai = get_cached_ai_result(url, content_hashes) if content_hashes else None

if cached_ai is not None:
    result = copy.deepcopy(cached_ai)
    result["url"] = url
    ai_status = "cached"
else:
    try:
        result = analyze_signals(url, signals)
        if content_hashes:
            store_ai_result(url, content_hashes, copy.deepcopy(result))
    except AIServiceError as e:
        result = _build_fallback_result(signals, str(e))

Three things to note:

Deep copy on both store and retrieve

The cache owns the dict it stores. The pipeline downstream mutates the result (enriches issues, ranks priorities, computes potential score). If we passed the same object around, the second hit would see the first hit's mutations. Deep copy is microseconds on a small report dict — irrelevant next to the HTTP call it saves.

Only cache successful AI results

If Gemini returns a 429 and we fall back to deterministic-only analysis, we don't cache that. The fallback is degraded; freezing it for six hours would be worse than paying for another call.

Always set url to the current request

Even within one domain, the user may have typedhttps://example.com vs https://www.example.com/. The report should show their URL, not whatever string was present on the first write.

TTL and eviction

Cache entries expire after 6 hours. Why 6 and not 24:

  • Too short and the cache doesn't help. Hourly monitors want the entry to survive overnight.
  • Too long and stale data lingers. If a site owner ships a fix, the next scan should reflect it within a reasonable window. Six hours is a tradeoff; twenty-four is embarrassing.

PROMPT_VERSION handles the long tail. A prompt bump invalidates everything cleanly even if entries would otherwise live longer.

We also cap the in-memory cache at 500 entries, oldest-first eviction. On a single worker that's enough for a moderate day of distinct content. Redis makes the cap and multi-worker sharing someone else's problem.

The result

Before: every scan called Gemini. Median wall time in the 25 to 30 second range when the model path dominated.

After: the first scan of a given site + content calls Gemini. Repeat scans of unchanged content return from cache in on the order of ~200 ms for the AI step — roughly a 150× reduction vs a full model round-trip on those hits (from our ai_ms logging on cache hits vs misses).

On Gemini's free tier, that also means we stop burning the daily quota on no-op re-analysis. An example load — 20 monitored sites, hourly checks — that would have been ~480 model calls a day becomes closer to one call per site per day of real content change, not one call per tick of the clock.

The cost savings aren't the interesting part. The interesting part is that we can offer continuous monitoring without the AI bill scaling with scan frequency. The bill scales with how often sites actually change.

What we'd do differently

We should have built this when monitoring shipped — not after.

The first version of GazeRank scanned one URL at a time, on demand. A content-hash cache would have been over-engineering. The moment scans became scheduled and automatic, the cache stopped being an optimization and became a requirement. We shipped monitoring without it, and the first week of runs cost a full Gemini call per site per hour. We only noticed when we started logging ai_ms per scan.

The lesson isn't “add caching early.” It's: the moment a request becomes recurring, ask what state you can carry between runs. Cron is supposed to be cheap. If every run costs as much as the first, you're not monitoring — you're just repeating the expensive path on a timer.

What's next

Two things we're considering:

Persistent cache (Redis)

Today the cache is in-memory, per worker. One instance is fine. Multiple workers mean a user can hit on one request and miss on the next if the load balancer switches nodes. Redis fixes that and survives restarts.

Smarter invalidation

If any single page changes, the whole scan misses. For a 12-page crawl where 11 pages are unchanged, we still pay for a full AI pass. Ideal would be page-level cache — but the system prompt is sitewide (duplicate titles, internal links, aggregate patterns), so per-page caching would break the analysis. Not worth it until per-scan cost is a real bottleneck again.

For now the tradeoff is right: hit the cache when nothing changed; when something did, re-analyze properly.

Run your first scan

Audit your site's AI search readiness

GazeRank runs dozens of deterministic checks plus AI analysis on any public URL. Scans of unchanged sites return far faster thanks to the caching above — and monitoring no longer means paying for the same answer 24 times a day.

Scan my site →