x-tweet-fetcher

August 8, 2026 · View on GitHub

x-tweet-fetcher

Fetch X/Twitter tweets, replies, timelines, lists, and articles — no login, no API keys.

License: MIT Python 3.10+ GitHub stars

Three backends · Auto fallback · Unified JSON schema · Built for AI agents

Quick Start · Backends · Capabilities · Python API · Self-hosted Nitter · Migrating from v1


😤 Problem

You: fetch that tweet / list / article for me
AI:  I can't access X/Twitter. Please copy-paste the content manually.

You: ...seriously?

X has no free API. Scraping gets you blocked. Browser automation is fragile in headless environments.

x-tweet-fetcher solves this with smart backend routing: FxTwitter for single tweets (zero deps), Nitter for timelines and search (direct HTTP), a browser driver for everything else — with automatic fallback between them.

🚀 Quick Start

git clone https://github.com/ythx-101/x-tweet-fetcher
cd x-tweet-fetcher && pip install .

# Single tweet — works instantly, zero configuration
xtf --url https://x.com/user/status/1234567890

# User timeline (needs a Nitter instance, see below)
export XTF_NITTER=http://127.0.0.1:8788
xtf --user elonmusk --limit 20

# Search
xtf --search "openclaw" --limit 10

# Human-readable output instead of JSON
xtf --user elonmusk --text-only

Prefer not to install? python3 scripts/fetch_tweet.py --url ... works straight from the clone (same flags).

🔀 Three Backends

BackendDepsSpeedCovers
fxtwitterNone (stdlib)⚡⚡Single tweets, user profiles
nitterA Nitter instanceTimeline, search, replies, mentions
browserCamofox or Playwright🐢Everything above + Lists + X Articles
auto (default)Best available⚡→🐢Nitter first, browser fallback
xtf --user elonmusk                    # auto (default)
xtf --user elonmusk --backend nitter   # direct HTTP only
xtf --list 1455045069516357634         # lists always use the browser

Browser driver defaults to Camofox (localhost:9377). Playwright users:

pip install ".[playwright]"            # from the clone
export XTF_BROWSER=playwright          # or: --browser-driver playwright

📊 Capabilities

FeatureFlagBackend
Single tweet (text, stats, media, quotes)--urlfxtwitter
Reply comments (threaded)--url --repliesnitter / browser
User timeline (paginated)--usernitter / browser
Search--searchnitter
User profile--user-infofxtwitter → nitter
X List tweets--listbrowser
X Article full text--articlebrowser
Mentions monitor (incremental, cron-friendly)--monitornitter / browser
Archive fetch results (dedupe, SQLite)--ledger <db>any
Search / stats the archive (offline)--ledger <db> --query/--statsoffline

Exit codes (cron-friendly): 0 success / no new mentions · 1 error / new mentions found · 2 monitor setup error.

Errors are machine-readable. Every failure carries error (human message) plus error_code — one of invalid_input, not_found, rate_limited, upstream_down, backend_unavailable, all_backends_failed — so agents can branch on it. all_backends_failed additionally includes per-backend error_causes.

📚 推文库 (Ledger)

--ledger <db> turns xtf into a fetch + archive + query local tweet library: every timeline / search / list / replies / single-tweet fetch is archived into a SQLite DB, deduped by tweet_id (INSERT OR IGNORE, idempotent). Schema is compatible with the tweet-ledger (OpenClaw) tweets table, so the same DB can be read by both tools.

# Fetch + archive a timeline
xtf --user YuLin807 --limit 20 --ledger ~/tweets.db

# Search the archive (offline)
xtf --ledger ~/tweets.db --query "sop"

# Stats: totals, languages, media/urls, time ranges
xtf --ledger ~/tweets.db --stats

Behavior without --ledger is unchanged (3.0.0-compatible). Archiving never breaks a successful fetch — on failure the JSON envelope carries ledger_error instead. Single-tweet (fxtwitter) dicts lack tweet_id, so the CLI injects it from the URL; --replies results are archived with is_reply=1 and in_reply_to_status_id pointing at the parent tweet.

tweets table: tweet_id (PK) · created_at · full_text · lang · source_file · is_reply · in_reply_to_status_id · retweeted_status_id · quoted_status_id · urls_json · media_json · raw_json · imported_at

End-to-end integration test record: docs/e2e-integration.md.

🐍 Python API

from xtf import Router, NotFound, RateLimited

router = Router()                                  # backend="auto"
tweet   = router.fetch_tweet("user", "1234567890") # dict, v1-compatible shape
tweets  = router.fetch_timeline("user", limit=20)  # list[Tweet]
replies = router.fetch_replies("user", "1234567890")
results = router.search("openclaw", limit=10)

for tw in tweets:
    print(tw.author, tw.likes, tw.text)
    print(tw.to_dict())                            # JSON-ready

All backends normalize into one Tweet / Reply / Profile / Article schema — your downstream prompt only ever needs to describe one shape.

⚙️ Configuration

Everything is an environment variable (CLI flags override):

VariableDefaultMeaning
XTF_NITTERhttp://127.0.0.1:8788Comma-separated Nitter instances, tried in order with failover
XTF_BROWSERcamofoxBrowser driver: camofox or playwright
XTF_BROWSER_PORT9377Camofox HTTP port
XTF_LANGzhMessage language: zh or en
XTF_CACHE_DIR~/.x-tweet-fetcherMentions-monitor cache

NITTER_URL (the v1 name) is still honored as a fallback for XTF_NITTER.

🏗 Self-hosted Nitter

Public Nitter instances are unreliable and frequently dead. Self-hosting is strongly recommended for timeline/search/replies:

# See https://github.com/zedeus/nitter for full setup
docker run -d -p 8788:8080 --name nitter zedeus/nitter:latest
export XTF_NITTER=http://127.0.0.1:8788

Multiple instances failover automatically:

export XTF_NITTER=http://127.0.0.1:8788,https://your-backup-instance.example

If no instance is reachable, you get a clear error (error_code: "all_backends_failed", with each backend's reason — e.g. backend_unavailable — under error_causes) telling you exactly what to set. Never a silent empty result.

📁 Project Structure

src/xtf/
├── models.py        # Tweet / Reply / Profile / Article dataclasses
├── backends/
│   ├── fxtwitter.py # single tweets + profiles
│   ├── nitter.py    # direct HTTP, multi-instance failover
│   └── browser.py   # Camofox / Playwright snapshot fetching
├── parsers/         # pure functions, locked by fixture tests
├── router.py        # auto-fallback chain
├── monitor.py       # incremental mentions monitor
└── cli.py           # the `xtf` command
scripts/fetch_tweet.py   # v1-compatible entry point (thin shim)
tests/fixtures/          # captured page structures — regression protection

🔄 Migrating from v1

python3 scripts/fetch_tweet.py still works with all v1 flags and exit codes, and JSON fields are unchanged for every mode except --search, whose per-tweet schema is now unified with --user (fields renamed, url/has_media/media_urls dropped). See MIGRATION.md for the full list, including where the analytics/China/Obsidian scripts went (spoiler: their own repos — this project is now purely about fetching tweets; the old world lives at the v1-legacy tag).

🧪 Development

pip install -e ".[dev]"
pytest          # all parsers locked by fixture tests
ruff check src tests

When Nitter or X change their page structure, capture a fresh snapshot into tests/fixtures/ — the failing test will show exactly which parser and field broke.

🙏 Acknowledgments

  • Nitter by zedeus — self-hosted Twitter frontend
  • FxTwitter — public API for single tweet data
  • Camofox — anti-fingerprint browser, default browser driver
  • Playwright — alternative browser automation driver
  • OpenClaw — AI agent framework this tool grew up in

📜 License

MIT


Three backends. Auto fallback. Built for AI agents.

GitHub · Issues · #22 Teahouse · Agent Waystation