Feed screenshots to GPT-4V, Claude, and Gemini — preprocessing that saves tokens

You want to build a tool that looks at a webpage and answers questions about it. Product research, competitive analysis, accessibility auditing, UI regression detection — the use cases are endless.

The obvious approach: screenshot the page, hand the image to GPT-4 Vision (or Claude, or Gemini), ask your question in the prompt. Works great on cheerful demo URLs. Breaks the moment you point it at anything real.

This post covers the four preprocessing steps that turn a naive "screenshot + prompt" pipeline into something you can charge for.

Why raw screenshots break AI vision

Three problems compound:

  1. Cookie banners and ad overlays. These cover the interesting content. The model dutifully describes the banner instead of the page.
  2. Tall pages. Real websites are 5,000–30,000 pixels tall. GPT-4V and Claude cap at 8,000 px in the long dimension; anything longer gets silently downscaled, losing the text you cared about.
  3. Cost. A full-page screenshot at 1920×12000 costs ~4,000 vision tokens (~$0.04/call on GPT-4V) even before your text prompt. That adds up when you're doing hundreds of URLs.

Every one of those problems is fixable with the right screenshot options. Here's the pipeline.

Step 1: strip banners, ads, and popups

The AI doesn't care about your GDPR banner. Turn them off at capture time:

const { url: cleanScreenshot } = await snap.screenshot({
  url: "https://competitor.com/pricing",
  format: "png",
  full_page: true,
  block_ads: true,          // strips known ad networks
  remove_popups: true,      // dismisses cookie / modal overlays
  click_accept: true,       // auto-clicks Accept on OneTrust, Cookiebot, etc.
  press_escape: true,       // fallback for modals without a close button
  hide_selectors: [         // belt-and-suspenders: hide known chat widgets
    ".intercom-launcher",
    "#hubspot-messages-iframe-container",
    ".drift-widget",
  ],
});

The resulting screenshot is what a real user would see three seconds after arriving on the page — not the first-frame mess.

Step 2: slice tall pages under the 8000 px cap

Both GPT-4V and Claude Vision reject images with a dimension over ~8000 pixels (they downscale, which destroys legibility). Rather than fighting the limit, slice the page yourself:

const result = await snap.screenshot({
  url: "https://competitor.com/pricing",
  format: "png",
  full_page: true,
  block_ads: true,
  remove_popups: true,
  full_page_slices: true,                     // enable slicing
  full_page_slice_height: 4000,               // safe under 8k limit
  full_page_slice_overlap_height: 200,        // preserve continuity
});

console.log(result.url);       // the full stitched image
console.log(result.slices);
// [
//   { index: 0, offset_y: 0,    width: 1920, height: 4000, url: "..." },
//   { index: 1, offset_y: 3800, width: 1920, height: 4000, url: "..." },
//   { index: 2, offset_y: 7600, width: 1920, height: 3200, url: "..." },
// ]

You now have 3 images, each safely under the model's dimension cap. Feed them in sequence to the vision model with a system prompt that explains they're pages of a longer capture.

Notice the 200-pixel overlap: it prevents a line of text from getting split exactly at a slice boundary. Cheap insurance.

Step 3: send only what you need

Now the token math. GPT-4V charges based on image size:

Image dimensionsApprox. GPT-4V tokensCost / call
1920 x 12000 (full page)~4000~$0.040
3 slices at 1920 x 4000 ~2200 (770/slice x 3)~$0.022
1 slice at 1920 x 4000 ~770 ~$0.008

If your question only needs the top of the page ("what's the hero headline?"), send only slice 0. That's a 5× cost reduction versus sending the whole page.

Combine with the metadata extractor for even more savings:

const result = await snap.screenshot({
  url: "https://competitor.com/pricing",
  format: "png",
  full_page: false,           // just the viewport
  extract_metadata: true,     // no extra token cost - text-based
});

console.log(result.extracted_metadata);
// {
//   title: "Pricing - Competitor",
//   favicon: "https://competitor.com/favicon.ico",
//   open_graph: {
//     "og:title": "Competitor Pricing",
//     "og:description": "Simple, transparent pricing...",
//   },
//   fonts: ["Inter", "IBM Plex Mono"],
//   http_status: 200,
// }

Feed the metadata as plain text alongside just the viewport screenshot. The model gets structured context for free (metadata is a text extraction, not a vision call).

Step 4: give the model the DOM if pictures aren't enough

Some questions are answered better by markup than by pixels. "Does this page have an SEO description?" is a text query; "does this page have a red button in the hero?" is a vision query. Route accordingly.

const result = await snap.screenshot({
  url: "https://competitor.com/pricing",
  format: "png",
  full_page: true,
  extract_text: true,   // stripped rendered text (post-JS)
  extract_html: true,   // full DOM (post-JS)
});

// For text queries: feed only extract_text + one thumbnail
// For layout queries: feed slices + optionally extract_html

Text extraction runs after JavaScript has finished, so single-page apps with rendered content still work. Same billing (1 credit) as a normal capture; the text is a bonus in the response.

End-to-end example: competitor pricing analysis

Here's a realistic pipeline that costs ~$0.02 per URL total (capture + vision + text output).

import GetSnap from "getsnap";
import OpenAI from "openai";

const snap = new GetSnap(process.env.GETSNAP_KEY!);
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY! });

async function analyzeCompetitor(url: string) {
  // 1. Capture with all the anti-noise defaults on
  const capture = await snap.screenshot({
    url,
    format: "png",
    full_page: true,
    block_ads: true,
    remove_popups: true,
    click_accept: true,
    full_page_slices: true,
    full_page_slice_height: 4000,
    full_page_slice_overlap_height: 200,
    extract_metadata: true,
  });

  // 2. Feed the metadata + first two slices to GPT-4V
  const response = await openai.chat.completions.create({
    model: "gpt-4-vision-preview",
    messages: [{
      role: "user",
      content: [
        { type: "text", text:
          `You're analyzing a competitor's pricing page. Return JSON:\n` +
          `{ "plans": [{ "name": "...", "price_usd": 0, "features": [] }] }\n\n` +
          `Page title: ${capture.extracted_metadata?.title}\n` +
          `Meta description: ${capture.extracted_metadata?.open_graph?.["og:description"] ?? ""}`,
        },
        ...capture.slices!.slice(0, 2).map((s) => ({
          type: "image_url" as const,
          image_url: { url: s.url, detail: "high" as const },
        })),
      ],
    }],
    response_format: { type: "json_object" },
  });

  return JSON.parse(response.choices[0].message.content!);
}

// Run against a competitor set
for (const url of ["https://compA.com/pricing", "https://compB.com/pricing"]) {
  console.log(url, await analyzeCompetitor(url));
}

Notes:

Python version

import os
from getsnap import GetSnap
from openai import OpenAI

snap = GetSnap(api_key=os.environ["GETSNAP_KEY"])
client = OpenAI()

def analyze_competitor(url: str) -> dict:
    capture = snap.screenshot(
        url=url,
        format="png",
        full_page=True,
        block_ads=True,
        remove_popups=True,
        click_accept=True,
        full_page_slices=True,
        full_page_slice_height=4000,
        full_page_slice_overlap_height=200,
        extract_metadata=True,
    )
    meta = capture.get("extracted_metadata", {}) or {}

    response = client.chat.completions.create(
        model="gpt-4-vision-preview",
        messages=[{
            "role": "user",
            "content": [
                {"type": "text", "text":
                    f"Analyze this competitor's pricing page. Return JSON: "
                    f'{{"plans":[{{"name":"...","price_usd":0,"features":[]}}]}}'
                    f"\n\nTitle: {meta.get('title','')}\n"
                    f"Description: {meta.get('open_graph',{}).get('og:description','')}"
                },
                *[
                    {"type": "image_url",
                     "image_url": {"url": s["url"], "detail": "high"}}
                    for s in capture["slices"][:2]
                ],
            ],
        }],
        response_format={"type": "json_object"},
    )
    return response.choices[0].message.content

Claude Vision + Gemini equivalents

The preprocessing is model-agnostic. The only thing that changes is how you attach the images:

Claude 3.5 Sonnet Vision

const response = await anthropic.messages.create({
  model: "claude-3-5-sonnet-20241022",
  max_tokens: 1024,
  messages: [{
    role: "user",
    content: [
      { type: "text", text: "What's the pricing model?" },
      ...capture.slices!.slice(0, 2).map((s) => ({
        type: "image" as const,
        source: { type: "url" as const, url: s.url },
      })),
    ],
  }],
});

Gemini 1.5 Pro

const response = await gemini.generateContent({
  contents: [{
    role: "user",
    parts: [
      { text: "What's the pricing model?" },
      ...capture.slices!.slice(0, 2).map((s) => ({
        fileData: { fileUri: s.url, mimeType: "image/png" },
      })),
    ],
  }],
});

All three model families accept image URLs directly (no need to base64-encode), so the getsnap.dev CDN URLs go straight in.

Caching pays for itself

A subtle but real cost saving: getsnap.dev caches identical requests. If you re-analyze the same competitor page a week later with the same options, you get cached: true back — 0 credits, 12 ms response time. The AI vision call is what actually costs you money at that point.

Slice URLs are stable across cache hits too, so if you keep the slice URLs in your DB and re-feed them to the model, you skip both the screenshot cost AND the CDN download cost on the model side.

Build your AI vision pipeline on solid foundations

Free tier: 100 captures/month. Full-page slicing, metadata extraction, and all anti-noise flags included from day one.

Get free API key

Recap

A well-preprocessed screenshot is worth 5× a naive one. Same model, same prompt, 40–80% lower token bill.

Related reading