SERP API 6 min read

Build a SERP Rank Tracker with Reader Page Snapshots

Build a Python SERP rank tracker that saves organic positions and Reader Markdown snapshots in SQLite, using data.organic and exact-domain matching.

(Updated: ) • 1,075 words

A rank tracker records where a domain appears for a defined query, country, language, and date. This tutorial adds a second artifact: a Reader Markdown snapshot of the first organic page. Keeping rank observations and page content in separate SQLite tables makes later manual review possible without treating a single ranking movement as proof of why it happened.

What this tutorial owns

Use this workflow when you want to compare an observed organic position with the content of a leading result. For a general tracking schema, retry policy, and dashboard design, use the rank-tracking API workflow. For a small keyword-to-CSV script, use the Python CSV tracker. This page focuses on preserving a page snapshot alongside each check.

What the APIs actually return

Send a JSON POST to https://www.searchcans.com/api/v1/search with s for the query and t: "google". In the response verified on 2026-09-29, data is an object and data.organic is the result array. Each organic result has a link and position. Iterating over data itself would iterate object keys, not search results.

The Reader endpoint is POST https://www.searchcans.com/api/v1/url. Send a result URL as s, set t: "url", and read the returned data.markdown. A page may fail extraction or produce little text, so the rank row should be saved before attempting the snapshot. See the Reader API documentation for current extraction options.

Build a small SQLite evidence archive

Install Python 3.9 or newer and requests, then set the SEARCHCANS_API_KEY environment variable. The example creates rank_checks for observed positions and page_snapshots for extracted Markdown. It requests the first search-results page; a missing domain means “not present in the returned organic results,” not “outside Google’s top 100.”

import os
import sqlite3
from datetime import datetime, timezone
from urllib.parse import urlsplit
import requests

API_KEY = os.environ["SEARCHCANS_API_KEY"]
SEARCH_URL = "https://www.searchcans.com/api/v1/search"
READER_URL = "https://www.searchcans.com/api/v1/url"
DB_FILE = "rank_tracking.db"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}

def setup_database():
    with sqlite3.connect(DB_FILE) as db:
        db.execute("""CREATE TABLE IF NOT EXISTS rank_checks (
            checked_at TEXT NOT NULL,
            keyword TEXT NOT NULL,
            target_domain TEXT NOT NULL,
            position INTEGER,
            ranking_url TEXT,
            top_url TEXT
        )""")
        db.execute("""CREATE TABLE IF NOT EXISTS page_snapshots (
            checked_at TEXT NOT NULL,
            source_url TEXT NOT NULL,
            markdown_content TEXT NOT NULL
        )""")

def api_post(url, payload):
    response = requests.post(url, json=payload, headers=HEADERS, timeout=45)
    response.raise_for_status()
    result = response.json()
    if result.get("code") != 0:
        raise RuntimeError(result.get("msg") or "API request failed")
    data = result.get("data")
    if not isinstance(data, dict):
        raise ValueError("Expected a data object")
    return data

def same_domain(link, target_domain):
    host = (urlsplit(link).hostname or "").lower().removeprefix("www.")
    target = target_domain.lower().removeprefix("www.").strip(".")
    return host == target or host.endswith("." + target)

def record_keyword(keyword, target_domain):
    data = api_post(SEARCH_URL, {"s": keyword, "t": "google", "p": 1,
                                 "country": "us", "language": "en"})
    organic = data.get("organic")
    if not isinstance(organic, list):
        raise ValueError("Expected data.organic to be a list")
    checked_at = datetime.now(timezone.utc).isoformat()
    top_url = organic[0].get("link") if organic and isinstance(organic[0], dict) else None
    position, ranking_url = None, None
    for index, item in enumerate(organic, start=1):
        if not isinstance(item, dict):
            continue
        link = item.get("link") or ""
        if same_domain(link, target_domain):
            raw_position = item.get("position")
            position = raw_position if isinstance(raw_position, int) and raw_position > 0 else index
            ranking_url = link
            break
    with sqlite3.connect(DB_FILE) as db:
        db.execute("INSERT INTO rank_checks VALUES (?, ?, ?, ?, ?, ?)",
                   (checked_at, keyword, target_domain, position, ranking_url, top_url))
    return checked_at, top_url

def record_page_snapshot(checked_at, top_url):
    if not top_url:
        return
    data = api_post(READER_URL, {"s": top_url, "t": "url", "mode": 1})
    markdown = data.get("markdown")
    if not isinstance(markdown, str) or not markdown.strip():
        raise ValueError("Reader returned no Markdown")
    with sqlite3.connect(DB_FILE) as db:
        db.execute("INSERT INTO page_snapshots VALUES (?, ?, ?)",
                   (checked_at, top_url, markdown))

if __name__ == "__main__":
    setup_database()
    for keyword, domain in [("python web scraping best practices", "realpython.com"),
                            ("AI agent framework", "github.com")]:
        try:
            checked_at, top_url = record_keyword(keyword, domain)
            record_page_snapshot(checked_at, top_url)
        except (requests.RequestException, RuntimeError, ValueError) as error:
            print(f"{keyword}: {error}")

Run the script once with two test queries, then inspect the tables before scheduling it. The example deliberately makes one Reader request for the first organic page per successful search. If you only need rank history, omit that call. Add bounded retries, logging, and a retention policy before production use; a Reader failure should not erase an already stored rank observation.

Interpret results without overclaiming

The match compares parsed hostnames, so notgithub.com does not match github.com; subdomains do. The code records the API’s position when present and falls back to the order of returned organic items. It does not infer impressions, clicks, a local-city rank, or the cause of a ranking change from one SERP snapshot.

For 1,000 keywords checked once daily over 30 days, the search workload is 30,000 requests before retries and extra pages. Reader extraction adds a separate request and credit cost for every page captured. Estimate both against the current plans and credit rules. This article does not assume that self-hosting or any provider is a fixed percentage cheaper than another.

Questions before scaling

Q: Why store Markdown if the SERP already has a snippet?

A: A snippet is a short search-result excerpt. A Reader snapshot can preserve more of the page that appeared at check time, subject to extraction success and the page’s access rules. It supports review; it does not prove what caused a ranking change.

Q: Is a missing domain the same as rank zero?

A: No. The example stores NULL when the domain is absent from the returned organic items. Record the query, region, time, page number, and any API error separately so a missing observation is not mistaken for a verified ranking position.

Q: What should be added before a larger rollout?

A: Add bounded retries for transient failures, concurrency aligned with your available Parallel Lanes, a per-query schedule, an error log, and a policy for how long to keep third-party page snapshots. Test on a small set first and compare the stored rows with the raw API response.

Tags:

SERP API Tutorial SEO Web Scraping Python API Development
SearchCans Team

SearchCans Team

SERP API & Reader API Experts

The SearchCans engineering team builds high-performance search APIs serving developers worldwide. We share practical tutorials, best practices, and insights on SERP data, web scraping, RAG pipelines, and AI integration.

Ready to build with SearchCans?

Test SERP API and Reader API with 100 free credits. No credit card required.