Anssol
Back to Skills
Development

Skill Seekers

Convert documentation sites, GitHub repos, and PDFs into ready-to-use Claude skills.

metadocsauthoring
<p align="center"> <img src="docs/assets/logo.png" alt="Skill Seekers" width="200"/> </p>

Skill Seekers

English | 简体中文 | 日本語 | 한국어 | Español | Français | Deutsch | Português | Türkçe | العربية | हिन्दी | Русский

Version License: MIT Python 3.10+ MCP Integration Tested PyPI version PyPI - Downloads Website GitHub Repo stars PyPI Downloads

<a href="https://trendshift.io/repositories/18329" target="_blank"><img src="https://trendshift.io/api/badge/repositories/18329" alt="yusufkaraaslan%2FSkill_Seekers | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a>

🧠 The data layer for AI systems. Skill Seekers turns documentation sites, GitHub repos, PDFs, videos, notebooks, wikis, and more — 18 source types — into structured knowledge assets, ready to power AI Skills (Claude, Gemini, OpenAI), RAG pipelines (LangChain, LlamaIndex, Pinecone), and AI coding assistants (Cursor, Windsurf, Cline). Prepare once, export to 22 targets.

💛 Sponsors

<!-- SPONSORS:START -->

Launch Partner

<p align="center"> <a href="https://www.atlascloud.ai/"><img src="docs/assets/sponsors/atlas-cloud.png" alt="Atlas Cloud" width="200"></a><br/><sub><b>Launch Partner</b></sub> </p>

Atlas Cloud — A full-modal, OpenAI-compatible AI inference platform. Skill Seekers supports it as a packaging/enhancement target via --target atlas with ATLAS_API_KEY.

Silver Sponsors

<p align="center"> <a href="https://www.rapidproxy.io/?utm_source=skillseekers&utm_medium=sponsor"><img src="docs/assets/sponsors/rapidproxy.png" alt="RapidProxy" width="140"></a><br/><sub><b>Sponsor — Silver</b></sub> </p> <!-- SPONSORS:END -->

Become a sponsor · GitHub Sponsors


🚀 Quick Start

# 1. Install
pip install skill-seekers

# 2. Create a skill from any source
skill-seekers create https://docs.djangoproject.com/

# 3. Package it for your AI platform
skill-seekers package output/django --target claude

You now have output/django-claude.zip, ready to use.

# Pick a different AI agent for enhancement (default: claude)
skill-seekers create https://docs.djangoproject.com/ --agent kimi
skill-seekers create https://docs.djangoproject.com/ --agent-cmd "my-custom-agent run"

🛰️ AI-driven project scan

Point scan at a project and an AI agent reads its manifests, README, Dockerfile/CI and sampled source imports — then emits one config per detected framework, plus a <project>-codebase.json for your own code:

skill-seekers scan ./my-react-app --out ./configs/scanned/
# → react.json, vite.json, tailwind.json, jest.json, my-react-app-codebase.json

skill-seekers create ./configs/scanned/react.json

If a detection has no existing preset, the AI generates a fresh config; on exit you can optionally publish it back to the community registry.

All 18 source types

skill-seekers create facebook/react            # GitHub repository
skill-seekers create ./my-project              # Local codebase
skill-seekers create manual.pdf                # PDF
skill-seekers create report.docx               # Word
skill-seekers create book.epub                 # EPUB
skill-seekers create notebook.ipynb            # Jupyter
skill-seekers create openapi.yaml              # OpenAPI/Swagger
skill-seekers create presentation.pptx         # PowerPoint
skill-seekers create guide.adoc                # AsciiDoc
skill-seekers create page.html                 # Local HTML (or a whole dir)
skill-seekers create feed.rss                  # RSS/Atom
skill-seekers create curl.1                    # Man page

# Video (YouTube, Vimeo, or local — needs skill-seekers[video])
skill-seekers create --video-url https://www.youtube.com/watch?v=... --name mytutorial
skill-seekers create --setup                   # auto-install GPU-aware visual deps

skill-seekers create --space-key TEAM --name wiki               # Confluence
skill-seekers create --database-id ... --name docs              # Notion
skill-seekers create --chat-export-path ./slack-export --name team-chat  # Slack/Discord

See the Scraping Guide for every source type and its options.


📦 Installation

pip install skill-seekers              # Core: scraping, GitHub, PDF, packaging
pip install skill-seekers[all-llms]    # + every LLM platform
pip install skill-seekers[mcp]         # + MCP server
pip install skill-seekers[all]         # Everything

Not sure what you need? Run the wizard: skill-seekers-setup

<details> <summary><b>All installation extras</b></summary>
InstallAdds
skill-seekers[gemini]Google Gemini support
skill-seekers[openai]OpenAI ChatGPT support
skill-seekers[all-llms]All LLM platforms
skill-seekers[mcp]MCP server for Claude Code, Cursor, etc.
skill-seekers[video]YouTube/Vimeo transcript & metadata extraction
skill-seekers[video-full]+ Whisper transcription & visual frame extraction
skill-seekers[jupyter]Jupyter Notebook support
skill-seekers[pptx]PowerPoint support
skill-seekers[confluence]Confluence wiki support
skill-seekers[notion]Notion pages support
skill-seekers[rss]RSS/Atom feed support
skill-seekers[chat]Slack/Discord chat export support
skill-seekers[asciidoc]AsciiDoc support
skill-seekers[all]Everything

Video visual deps (GPU-aware): after installing skill-seekers[video-full], run skill-seekers create --setup to auto-detect your GPU and install the matching PyTorch variant + easyocr.

</details>

Prerequisites: Python 3.10+, Git. New here? → Bulletproof Quick Start 🎯


📚 Documentation

I want to...Read this
Get started quicklyQuick Start — 3 commands to your first skill
Understand the conceptsCore Concepts
Scrape sourcesScraping Guide — all 18 source types
Enhance skills with AIEnhancement Guide · Enhancement Modes
Export skillsPackaging Guide
Build workflowsWorkflows
Look up a commandCLI Reference — all 19 commands
ConfigureConfig Format · Environment Variables
Set up MCPMCP Setup · MCP Reference
Integrate with RAG / IDEsLangChain · RAG Pipelines · Cursor · Windsurf · Cline
Handle huge doc setsLarge Documentation — 10K–40K+ pages
Understand the architectureUML Architecture — 14 diagrams
Fix a problemTroubleshooting

Complete documentation index: docs/README.md


🎯 What you get

Use caseOutputPowers
AI SkillsComprehensive SKILL.md + reference filesClaude Code, Gemini, GPT
RAG pipelinesChunked documents with rich metadataLangChain, LlamaIndex, Haystack
Vector databasesPre-formatted data ready for upsertPinecone, Chroma, Weaviate, FAISS, Qdrant
AI coding assistantsContext files your IDE AI reads automaticallyCursor, Windsurf, Cline, Continue.dev

Export targets (22)

skill-seekers package output/react --target claude      # → Claude Skill (ZIP + YAML)
skill-seekers package output/react --target langchain   # → LangChain Documents
skill-seekers package output/react --target llama-index # → LlamaIndex TextNodes
skill-seekers package output/react --target ibm-bob     # → IBM Bob skill directory

LLM platforms (12): claude · gemini · openai · minimax · opencode · kimi · deepseek · qwen · openrouter · together · fireworks · markdown RAG & vector (8): langchain · llama-index · haystack · chroma · faiss · weaviate · qdrant · pinecone Other (2): atlas · ibm-bob

See the Feature Matrix for per-platform support details.

Why it matters

  • 99% faster — days of manual data prep → 15–45 minutes
  • 🎯 Real skill quality — 500+ line SKILL.md files with examples, patterns, and guides
  • 📊 RAG-ready chunks — smart chunking preserves code blocks and context
  • 🔄 Multi-source — combine docs + GitHub + PDFs + videos into one knowledge asset
  • 🌐 One prep, every target — export to 22 targets without re-scraping
  • Battle-tested — 3,900+ tests, 68 workflow presets, production-ready

✨ Key capabilities

<details> <summary><b>Documentation scraping</b> — SPA discovery, llms.txt, smart categorization</summary>

Three-layer discovery for JavaScript SPA sites (sitemap.xmlllms.txt → headless browser rendering), automatic llms.txt detection (10× faster when present), smart topic categorization, and a lenient HTML parser fallback so broken markup still scrapes.

Scraping Guide · llms.txt Support

</details> <details> <summary><b>GitHub & codebase analysis (C3.x)</b> — AST parsing, pattern detection, how-to guides</summary>

Three-stream architecture: code analysis (AST, design patterns, tests), documentation (README, docs/, wiki), and community (issues, PRs, metadata). The C3.x pipeline adds 10 GoF pattern detectors across 9 languages, usage examples extracted from tests, AI-written how-to guides, config extraction, and architecture overviews.

skill-seekers create ./my-project --preset quick          # 1–2 min, surface level
skill-seekers create ./my-project --preset standard       # balanced (default)
skill-seekers create ./my-project --preset comprehensive  # deep, exhaustive

Pattern Detection · How-To Guides · Test Example Extraction

</details> <details> <summary><b>AI enhancement</b> — API or local agents, 68 workflow presets</summary>

Every AI call runs through one transport, in API mode (Anthropic, Google Gemini, OpenAI, Moonshot/Kimi, MiniMax) or LOCAL mode (Claude Code, Kimi Code, Codex, Copilot, OpenCode, custom agents — no API costs). Control depth with --enhance-level 0-3 and pick an agent with --agent.

Enhancement Guide · Enhancement Modes · Multi-Agent Setup

</details> <details> <summary><b>Unified multi-source scraping</b> — combine many sources into one skill</summary>

One config can pull documentation, GitHub, PDFs, videos, and more into a single knowledge asset, with conflict detection and pairwise synthesis across sources.

Unified Scraping

</details> <details> <summary><b>Video extraction</b> — transcripts, frames, on-screen code</summary>

YouTube, Vimeo, and local files. Three-tier transcript fallback (subtitles → YouTube transcript API → local Whisper), plus optional visual extraction that OCRs on-screen code from sampled frames.

Video Guide

</details> <details> <summary><b>Quality, sync & scale</b></summary>

Quality scoring with a gate (skill-seekers quality output/react/ --threshold 7), provisional English readability metrics (informational — they never affect the score), doc-change detection with scheduled re-scrapes and notifications, streaming ingestion for very large doc sets, and incremental updates.

Large Documentation · Code Quality

</details>

🔌 MCP Integration (40 tools)

Skill Seekers ships an MCP server for Claude Code, Cursor, Windsurf, VS Code + Cline, and IntelliJ IDEA.

# stdio mode (Claude Code, VS Code + Cline)
python -m skill_seekers.mcp.server_fastmcp

# HTTP mode (Cursor, Windsurf, IntelliJ)
python -m skill_seekers.mcp.server_fastmcp --transport http --port 8765

Then just ask your assistant: "Package and upload the React skill."

MCP Setup · MCP Reference · HTTP Transport


🤖 Installing to AI agents

Skills install automatically into 19 AI coding agents:

skill-seekers install-agent output/react/ --agent cursor
skill-seekers install-agent output/react/ --agent all      # every detected agent
skill-seekers install-agent output/react/ --agent cursor --dry-run
AgentPathScope
Claude Code~/.claude/skills/Global
Cursor.cursor/skills/Project
VS Code / Copilot.github/skills/Project
Amp~/.amp/skills/Global
Goose~/.config/goose/skills/Global
OpenCode~/.opencode/skills/Global
Letta~/.letta/skills/Global
Aide~/.aide/skills/Global
Windsurf~/.windsurf/skills/Global
Neovate~/.neovate/skills/Global
Roo Code.roo/skills/Project
Cline.cline/skills/Project
Aider~/.aider/skills/Global
Bolt.bolt/skills/Project
Kilo Code.kilo/skills/Project
Continue~/.continue/skills/Global
Kimi Code~/.kimi/skills/Global
IBM Bob.bob/skills/Project

Uploading to Claude

export ANTHROPIC_API_KEY=sk-ant-...
skill-seekers package output/react/ --upload   # package + upload
skill-seekers upload output/react.zip          # upload an existing zip

No API key? Package it and upload output/react.zip manually at claude.ai/skills.

Upload Guide


⚙️ How it works

graph LR
    A[Documentation Website] --> B[Skill Seekers]
    B --> C[Scraper]
    B --> D[AI Enhancement]
    B --> E[Packager]
    C --> F[Organized References]
    D --> F
    F --> E
    E --> G[AI Skill .zip]
    G --> H[Upload to AI Platform]
  1. Scrape — extract every page (checking llms.txt first)
  2. Categorize — organize content into topics (API, guides, tutorials, …)
  3. Enhance — AI writes a comprehensive SKILL.md with examples
  4. Package — bundle into a platform-ready artifact
  5. Upload — ship it to your AI platform (optional)

Architecture

8 core modules + 5 utility modules (~200 classes):

ModulePurpose
CLICoreGit-style command dispatcher, source auto-detection
Scrapers18 source-type extractors on a shared build layer
Adaptors22 output platform formats behind one SkillAdaptor ABC
AnalysisC3.x codebase pipeline, 10 GoF pattern detectors
EnhancementAI improvement via a single AgentClient transport
PackagingPackage, upload, and install skills
MCPFastMCP server (40 tools, 10 tool modules)
SyncDoc change detection and notification

UML Architecture · API Reference · Skill Architecture


🆕 New in v3.9.0

  • HTML parser fallback for broken markup (#96) — severely malformed pages no longer scrape as empty; well-formed pages are byte-identical.
  • Transient-failure retries — the doc scraper (#97) and MCP fetch_config (#92) now retry connection blips and 5xx with backoff; 4xx still fails fast.
  • Whisper transcription fallback (#420) — local videos without subtitles finally get a real transcript.
  • MiniMax image OCR + registry-driven multimodal providers (#423) — providers declare their wire protocol and image capability; China-issued keys work against the right endpoint.
  • Token-lean GitHub issue defaults (#169) — GitHub skills no longer bundle full closed-issue history by default.
  • Env-driven CORS across all three servers (#422, #424) — no more wildcard origins with credentials.

Full history: CHANGELOG.md


📈 Performance

Documentation sizeTimeOutput
Small (< 100 pages)5–10 min~2 MB
Medium (100–500 pages)15–30 min~10 MB
Large (500–2,000 pages)30–60 min~40 MB
Huge (10K–40K+ pages)Use streamSee Large Documentation

🐛 Troubleshooting

skill-seekers doctor          # diagnose installation & environment
skill-seekers sync-config     # detect config drift

Common issues and fixes: Troubleshooting Guide · TROUBLESHOOTING.md


🤝 Contributing

Contributions are welcome — see CONTRIBUTING.md.


📝 License

MIT — see LICENSE.

🔒 Security

MseeP.ai Security Assessment Badge


🌐 Ecosystem

Skill Seekers is a multi-repo project:

RepositoryDescriptionLinks
Skill_SeekersCore CLI & MCP server (this repo)PyPI
skillseekerswebWebsite & documentationLive
skill-seekers-configsCommunity config repository
skill-seekers-actionGitHub Action for CI/CD
skill-seekers-pluginClaude Code plugin
homebrew-skill-seekersHomebrew tap for macOS

Want to contribute? The website and configs repos are great starting points for new contributors!

Want something like this for your team?

We build custom MCPs and Claude skills tailored to private workflows. Get in touch.

Start a Conversation

Ready to Ship Something?

Tell us about your project. We respond within one business day with honest scoping — not a sales pitch.

Get Started