sec02: EN soft sweep — +8 files (Fordham 18, von Franz Dreams 19, Hall 22, Meier 26, Whitmont/Perera 28); sec03/04 gaps re-verified (0 rescues); AGENTS.md: EN sweep procedure documented

This commit is contained in:
Dmitry Kokorin 2026-09-21 10:02:56 +03:00
parent 6334148b45
commit 1ff76b4270
23 changed files with 114 additions and 23 deletions

View file

@ -439,6 +439,44 @@ Special cases:
- ~4s between libgen requests, ~3s between RSL pages, ~2s between shop requests.
- All tools cache nothing by default; reruns are fine but avoid duplicate queries in a batch.
## EN original sweep (soft libgen queries) — validated 2026-09-20 (sec01: +50 files)
Exact full-title libgen queries miss a lot (indexed titles vary: reprints, e-books, variants).
The validated approach for closing the EN download gap of a section:
1. **Build the gap list** from `downloads/0N-*/MANIFEST.md` («EN full text unavailable /
not present» block) + SUMMARY Table 3 DL EN column (rows with «— (restricted)» / «—»).
2. **Soft libgen queries, 2–3 rounds** (author-centric, not title-guessing):
- round 1: `<author surname> <1–2 distinctive title words>`;
- round 2 (0-hit items): `<author surname>` alone, filter results by title/first name;
- round 3: alternate wording (shortened title, co-author, publisher word).
Whole-word matching still applies (full words, case-insensitive). `tools/lg.py` prints
`(unparsed…)` on 0-result pages — treat as 0 but confirm once with raw curl grep `edition.php?id=`.
3. **Collect candidates → batch metadata:** `https://libgen.vg/json.php?object=e&addkeys=*&ids=<comma-list>`
(≤10 ids) → `files` dict gives `f_id`+`md5` per file; then `object=f&addkeys=*&ids=<f_ids>`
for extension/filesize. Filter by year/edition to match the reading-list book (identity check:
title + author + era; e.g. SE15 vs SE22 look similar in truncated titles).
4. **Download queue (serial):** `tools/lgdl.py dl <f_id> <outdir> <name>`; resume works via `.part`
(curl `-C -`); the tool retries with fresh get.php keys. Huge files (>100 MB) can crawl at
~10 KB/s on libgen CDN (random stuck offsets) — let it stall-detect, or abandon in favor of
another edition of the same book (e.g. Great Mother: 643 MB Princeton 2015 abandoned →
64 MB Bollingen 1991).
5. **Monitor-kill trap (2026-09-20):** NEVER `pkill -f "lgdl.py"` (or any pattern that appears in
the queue bash's own command line) — it kills the queue itself. Kill only the worker by comm:
`ps -eo pid,comm,cmd | awk '$2=="python3" && /tools\/lgdl/ {print $1}' | xargs -r kill`.
6. **Verify after:** magic bytes (`%PDF` / `PK\x03\x04` for epub; discard HTML stubs) +
`pdfinfo` page count for small files — 2–14 pp «pdf» are fragments/essays, keep but mark
«фрагмент» in MANIFEST (e.g. Way of All Women 1933 = 2-pp, Self in Transformation = 14-pp JAP essay).
7. **archive.org (ia) as complement** for trade-restricted items: `tools/ia.py search "…"` →
open items only (`RESTRICTED` = borrow-only). Jung CW: single open item `CarlJungCollectedWorks`
(vols 1–18 + seminars, PDF+EPUB, direct `/download/<item>/<file>`).
8. **Docs to update after:** SUMMARY Table 3 DL EN + totals row, MANIFEST (EN section with lg f_id /
ia item per file; «not downloadable» list), per-card line «EN downloaded (date): <source> — <edition>»
(or «EN full text: NOT FOUND (date, 3 rounds)» for the true gaps), CONTEXT.md note. One commit per section.
Result sec01: 50 files from 58 missing (8 genuinely absent: 13, 16, 38, 42, 48, 50, 52, 54 +
643 MB 24a replaced by 1991 twin).
## Download phase conventions (user rule, 2026-07-18)
- **EN originals: libgen.vg FIRST** (the EN collection is bigger than the RU one — most trade