[{"data":1,"prerenderedAt":114},["ShallowReactive",2],{"seo-verification":3,"marketplace-app-en-archivebox":6},{"google":4,"bing":5},"EycwPY2XMyTkVzas3n1ygeNJFGAH513qrMjfDljzsMQ","",{"key":7,"data":8},"marketplace-app-en-archivebox",{"slug":9,"slugs":10,"categorySlugs":11,"name":16,"description":17,"phase":18,"unavailableReason":19,"docsUrl":20,"logo":21,"github":21,"tagline":22,"longDescription":23,"features":24,"useCases":33,"steps":43,"faq":59,"specs":75,"compatibleOs":90,"relatedApps":91,"relatedPosts":109,"category":111},"archivebox",{"fr":9,"en":9,"ar":9,"es":9},{"fr":12,"en":13,"ar":14,"es":15},"collaboration","collaboration-productivity","التعاون-والإنتاجية","colaboracion-productividad","ArchiveBox","Self-host your personal web archive on a VPS: save web pages, articles and media locally for offline reading, team sharing or regulatory compliance.",2,"awaiting_qa","https:\u002F\u002Fservorbit.com\u002Fblog\u002Fself-host-archivebox-web-archive-vps",null,"ArchiveBox — self-hosted web archive: save pages, videos and articles locally. MIT · 29 k+ stars.","ArchiveBox is an open-source (MIT) self-hosted web archiving tool (29 k+ GitHub stars) that turns any list of URLs into a fully browsable offline library. Each snapshot is saved in multiple formats simultaneously — HTML, PDF, screenshot PNG, WARC, DOM JSON, plaintext — so nothing is lost even if the source page disappears.\n\nInstalled on a ServOrbit VPS, ArchiveBox runs as a single Docker container (\u003C 400 MB RAM at rest) backed by a persistent data volume. The Django-based web interface lets you browse your archive, search full-text across all saved pages, replay media captures and share individual snapshots via signed URLs.\n\nArchiveBox also integrates yt-dlp for video archiving (YouTube, Vimeo, Twitter), wget for raw HTML snapshots, SingleFile for faithful full-page HTML, and Chromium for JavaScript-rendered pages. Scheduling is handled by a built-in cron system or by piping URLs from RSS feeds, bookmark exports or custom scripts via the CLI.\n\nArchiveBox is 100% self-hosted: your archive stays on your VPS, no cloud account is required and no data is ever sent to third parties.",[25,26,27,28,29,30,31,32],"Multi-format snapshots — each URL is saved as HTML, PDF, screenshot, WARC, DOM JSON and plaintext in a single operation.","yt-dlp integration — archive videos and audio from YouTube, Vimeo, Twitter and 1 000+ other sites alongside webpage snapshots.","Full-text search — Sonic or SQLite FTS indexes all saved page content, titles and annotations for instant recall.","Web UI — browse, replay, search and manage your archive through a clean Django-based interface.","CLI & automation — add URLs by batch, import from RSS feeds, bookmark exports (Pocket, Pinboard, Instapaper) or custom scripts.","Signed sharing links — share individual archive items with time-limited or permanent signed URLs without exposing your whole library.","Extensible extractors — toggle per-URL which extractors to run (wget, SingleFile, chrome, yt-dlp, git…) from the UI or `.env`.","MIT licence — no telemetry, no cloud account, no lock-in. Your data stays on your VPS.",[34,37,40],{"title":35,"body":36},"Compliance and legal archiving","Archive web-based evidence, regulatory documents and client-facing pages on your own infrastructure. ArchiveBox stores cryptographic hashes of each snapshot alongside the capture metadata, providing a tamper-evident chain of custody without relying on the Wayback Machine or a third-party SaaS.",{"title":38,"body":39},"Personal reading library","Save articles from newsletters, blogs and news sites before they go behind a paywall or vanish. ArchiveBox strips cookies and tracking, stores a clean HTML copy and a readable plaintext version, then indexes everything for full-text search — making your offline library searchable in seconds.",{"title":41,"body":42},"Research and journalism","Capture web sources at the moment of discovery with full-page screenshots, WARC files and metadata. Replay the exact state of a page as it appeared during your research — useful for fact-checking, sourcing citations and preserving evidence for investigative work.",[44,47,50,53,56],{"title":45,"body":46},"Order from the ServOrbit Marketplace","From your ServOrbit client area, install ArchiveBox in one click from the Marketplace: select the Collaboration category, choose ArchiveBox and confirm your order. The Docker container starts automatically on your VPS. You will receive a notification when the healthcheck responds, with the assigned port and the generated admin password.",{"title":48,"body":49},"Log in to the admin interface","Open an SSH tunnel to reach the interface without a domain: `ssh -L 5797:127.0.0.1:\u003Cport> root@\u003Cvps-ip>` then navigate to `http:\u002F\u002Flocalhost:5797\u002F`. Your admin credentials are in your ServOrbit client area under Applications → Reveal. Change the password immediately after your first login under Settings → Change password.",{"title":51,"body":52},"Archive your first URLs","Click Add in the top bar and enter one or more URLs (one per line). Select which extractors to run — HTML, PDF, screenshot and WARC are enabled by default. ArchiveBox queues the captures, downloads each snapshot in parallel and adds them to your archive index. Large batches can also be imported via the CLI: `archivebox add \u003C urls.txt`.",{"title":54,"body":55},"Schedule recurring captures","To archive URLs on a schedule, create a recurring cron job via the ArchiveBox CLI or a dedicated compose service (`archivebox schedule --every=day --depth=1 \u003Curl>`). RSS feeds and bookmark exports (Pocket JSON, Pinboard, Instapaper) can be imported automatically by piping them to `archivebox add`.",{"title":57,"body":58},"Attach a domain for permanent access","In your ServOrbit client area, attach a domain or subdomain to your ArchiveBox instance. ServOrbit configures nginx and provisions a Let's Encrypt certificate automatically. Once the domain is active, update the `ARCHIVEBOX_CSRF_TRUSTED_ORIGINS` environment variable in your compose file to match the new origin.",[60,63,66,69,72],{"q":61,"a":62},"Does ArchiveBox require a domain name?","No. ArchiveBox listens on the loopback interface and is reachable via SSH tunnel (`ssh -L 5797:127.0.0.1:\u003Cport> root@\u003Cvps-ip>`) without any domain. If you want a permanent HTTPS URL for your team, attach a subdomain from your ServOrbit client area or use the free VPS subdomain `{app}.{dns_slug}.servorbit-dns.com`.",{"q":64,"a":65},"Which formats does ArchiveBox save?","Each URL can produce: HTML snapshot (wget), full-page HTML (SingleFile), screenshot PNG, PDF, WARC, DOM JSON, plaintext, git clone (for code repos) and video\u002Faudio (yt-dlp for YouTube, Vimeo and 1 000+ sites). You can enable or disable individual extractors per-URL or globally in the `.env` file.",{"q":67,"a":68},"How much disk space does archiving use?","A typical article snapshot (HTML + screenshot + PDF) weighs 2–10 MB. Videos archived with yt-dlp can reach several gigabytes. Plan your VPS disk capacity accordingly — ServOrbit volumes can be expanded from your client area without interrupting the running container.",{"q":70,"a":71},"Can I share archived pages with others?","Yes. The admin interface lets you generate signed sharing links for individual snapshots. Recipients can replay the archived page in their browser without needing an ArchiveBox account. Public access to the whole archive can also be enabled via `PUBLIC_INDEX=True` in the environment.",{"q":73,"a":74},"How do I update ArchiveBox?","Pull the new image and recreate the container: `docker compose pull && docker compose up -d`. ArchiveBox runs Django migrations automatically on startup — your existing archive data in the named volume is preserved. Always check the release notes for breaking changes before updating.",{"ram":76,"cpu":77,"disk":78,"stack":79,"port":86,"license":87,"version":88,"os":89},"512 MB recommended","1 vCPU","20 GB+ (depends on archive volume)",[80,81,82,83,84,85],"Docker","Python 3.13","Django","SQLite \u002F PostgreSQL","Chromium","yt-dlp","5797","MIT","0.9.71","Ubuntu 24.04",[],[92,99,104],{"name":93,"slug":94,"categorySlug":13,"categoryName":95,"categoryColor":96,"logo":21,"tagline":97,"description":98},"Karakeep","karakeep","Collaboration & Productivity","text-indigo-400 bg-indigo-500\u002F10","Self-hosted AI bookmark manager — save anything, find everything. Auto-tags and full-text search powered by Meilisearch and Ollama or OpenAI.","AI-powered self-hosted bookmark manager: save links, notes, images and PDFs, let the AI auto-tag and summarize everything — with local Ollama or OpenAI. 3 Docker containers (app + Meilisearch + headless Chrome), port 3000, AGPL-3.0. 26 k ⭐ GitHub.",{"name":100,"slug":101,"categorySlug":13,"categoryName":95,"categoryColor":96,"logo":21,"tagline":102,"description":103},"FreshRSS","freshrss","Self-hosted multi-user RSS\u002FAtom reader — follow unlimited feeds, sync any mobile client via Google Reader API, refresh on your schedule. SQLite, single container.","Self-hosted multi-user RSS\u002FAtom aggregator that tracks your technical blogs, newsletters and news sources in a clean interface — synced to any device via the Google Reader and Fever APIs. SQLite, single container, under 128 MB RAM.",{"name":105,"slug":106,"categorySlug":13,"categoryName":95,"categoryColor":96,"logo":21,"tagline":107,"description":108},"Blinko","blinko","AI-powered self-hosted notes — capture ideas, search semantically, own your knowledge. GPL-3 · 11 k+ stars.","Open-source AI-powered note-taking: capture ideas as cards, find them by semantic search and keep your knowledge fully private on your VPS.",[110],"self-host-archivebox-vps",{"key":12,"slug":13,"name":95,"objective":112,"icon":113,"color":96},"Share information and collaborate.","collab",1790693365153]