README.md

July 27, 2026 · View on GitHub

FindJobs-Agent

Python 3.9+ License: MIT CI

English · 中文  ·  Quick Start · How It Works · Features


What is FindJobs-Agent?

A full-stack job search assistant that crawls postings from major tech companies, analyzes them with LLMs, parses your resume, and runs AI mock interviews — so you can focus on preparing, not sifting through job boards.

How It Works

Four pieces wired into one flow: a crawler pulls postings from company career sites, an LLM reads each one for requirements and skills, your resume gets parsed and scored against them, and any posting can drive an AI mock interview straight off its job description. The React frontend ties it together, so you go from "what's out there" to "let me practice for this one" without leaving the app.

FindJobs-Agent architecture

Features

  • Job crawler — pulls postings from Tencent, NetEase, ByteDance, Amazon and more, via API or Selenium, with automatic cleaning and normalization.
  • LLM analysis — extracts education/major requirements, scores skill tags (1–5), and classifies each posting into a job taxonomy.
  • Resume parsing & matching — parses PDF/Word resumes, scores skills, and computes a case-insensitive job-resume match percentage.
  • AI mock interview — generates questions from any job description and runs a multi-turn interview with real-time feedback.
  • SQLite persistence — analyzed postings are stored in a local jobs.db; existing CSV data is migrated automatically on first run, with CSV/JSON as fallback.

Project Structure

FindJobs-Agent/
├── FrontEnd/                # React frontend
│   ├── src/
│   │   ├── components/      # Page components
│   │   │   ├── JobsPage.tsx       # Job browsing
│   │   │   ├── ResumePage.tsx     # Resume analysis
│   │   │   └── InterviewPage.tsx  # AI interview
│   │   └── App.tsx
│   └── package.json
├── job_crawler_v2.py        # Multi-company crawler (primary)
├── job_crawler_selenium.py  # Selenium crawler
├── job_agent.py             # LLM job analysis agent
├── pipeline.py              # Data processing pipeline
├── api_server.py            # Flask API server
├── storage.py               # SQLite job store (jobs.db)
├── interview_agent.py       # AI interview module
├── resume_parser.py         # Resume parser
├── tag_rate.py              # Skill scoring
├── llm_client.py            # LLM client
├── tech_taxonomy.json       # Job taxonomy
├── all_labels.csv           # Skill tag library
└── requirements.txt

Quick Start

Prerequisites

  • Python 3.9+
  • Node.js 18+
  • Chrome (required for Selenium crawler)

1. Clone the repo

git clone https://github.com/he-yufeng/FindJobs-Agent.git
cd FindJobs-Agent

2. Install backend dependencies

pip install -r requirements.txt

3. Set up your API key

Create an API_key.md file with your OpenAI API key:

sk-your-api-key-here

4. Start the backend

python api_server.py

5. Start the frontend

cd FrontEnd
npm install
npm run dev

6. Open the app

Visit http://localhost:8080 in your browser.

Data Pipeline

pipeline.py chains crawl → analyze → score → serve. Run the whole thing, or a single stage:

python pipeline.py                                          # crawl + analyze + build site data
python job_crawler_v2.py -c tencent netease amazon -m 300   # crawl only (--list shows companies)
python pipeline.py --analyze-only --max-jobs 50             # analyze only (for testing)

API Endpoints

EndpointMethodDescription
/api/jobsGETList job postings
/api/jobs/<id>GETGet job details
/api/resume/uploadPOSTUpload resume
/api/resume/analyzePOSTAnalyze resume
/api/interview/startPOSTStart mock interview
/api/interview/answerPOSTSubmit interview answer

Roadmap

Crawl, analyze, resume match, and mock interview work end to end. The next steps widen the funnel and follow the hunt past the match:

  • More job sources — extend the crawler beyond the current company set to job boards and aggregators, so matching isn't limited to a fixed list.
  • Incremental crawls — track which postings were already seen and fetch only new ones, instead of re-crawling and re-analyzing the full set each run.
  • Application tracking — the backend is in: /api/applications keeps a per-job status board (bookmarked / applied / replied / interview / offer / rejected) in SQLite. The UI board on top of it is next.
  • Voice mock interviews — speech in and out for the AI interviewer, closer to a real screen than a text chat.

FindJobs-Agent is one of the applied agents I've built. A few others you might find useful:

  • CoreCoder — want to understand how a coding agent really works? Read the whole ~1k-line engine end to end, not a black box.
  • RepoWiki — dropped into an unfamiliar codebase? It gives you a guided wiki and a where-to-start reading path, a self-hostable DeepWiki alternative.
  • ContractGuard — catch the risky clauses before you sign: it reads contracts and flags the dangerous bits.
  • GitSense — want to contribute to open source? It finds issues worth your time and gauges whether your PR will get merged.
  • CodeABC — understand any codebase even if you don't code, built for non-programmers.

Contributing

Issues and pull requests are welcome!

License

MIT License