| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| archivebox-0.9.64-py3-none-any.whl | 2026-09-26 | 1.1 MB | |
| archivebox-0.9.64.tar.gz | 2026-09-26 | 959.5 kB | |
| COMMIT_SHA | 2026-09-26 | 41 Bytes | |
| SHA256SUMS | 2026-09-26 | 269 Bytes | |
| README.md | 2026-09-26 | 3.1 kB | |
| v0.9.64_ More compact UI on Snapshot detail and list pages source code.tar.gz | 2026-09-26 | 2.7 MB | |
| v0.9.64_ More compact UI on Snapshot detail and list pages source code.zip | 2026-09-26 | 3.1 MB | |
| Totals: 7 Items | 7.9 MB | 0 | |
Highlights
- 🗂️ Much more compact snapshot detail page header and list view
- 🗃️ Fixed migrated snapshots failing to open, including archives created in non-UTC time zones.
- 📄 ArchiveResult admin page improvements, better output files summary in admin views
- 📚 Improved article text extraction that preserves original image sizes, fixed SingleFile to avoid re-capturing Chrome's PDF viewer
- 🖼️ Improved capture of pages with images still loading, and made failed or empty downloads and WACZ exports easier to spot.
- 🔎 Fixed archive searches failing to start after reindexing.
Historic Context on the Recent Filesystem Migrations
0.7.x -> 0.9.x is the first time we've ever changed the filesystem layout of the data/ dir, because we're now moving to a model where snapshots are separated on disk by the user/{username} that owns them. This move also allowed us to solve longstanding issues with large collection performance that were caused by storing 10k+ directories in a single snapshot directory, the new layout puts them in a tree under user>recordtype>date>origin>uuid, which can hopefully be stable for a long time and scale to trillions of entries.
There are two ways we could've handle this (potentially massive) filesystem migration for people's existing data coming from 0.7.x:
- lazily move files to new locations when they are first accessed (to avoid needing to do a huge slow migration all at once)
- eagerly move files during the updating process, or ask users to run a command that does it all at once (
archivebox update --migrate-only)
Initially I tried to do it lazily on first access, but it is really hard to predict the load characterestics of the server / decide when to migrate things without wild spiky performance issues. e.g. if you load a big list of snapshots and open them all at once, it can easily crash the server unless we build a whole system to manage load. After much tinkering I decided it was better to have users explicitly run a command to migrate everything at once, and to build a few branches into the UI code to check for both the old structure and the new structure temporarily until migrations finish.
Predictably this has led to some bugs and in increased complexity around URLs and deeplinks into specific snapshot content (e.g. you may have noticed the icons from the snapshot list view lead to 404s when clicked on old migrated data). I apologize for these issues if you've run into them, but trust that your data is still there, it's just a bug in the UI code deciding to check for old locations vs new locations. Please help report bugs and open Github issues if you encounter anything like that!