sec02: EN soft sweep — +8 files (Fordham 18, von Franz Dreams 19, Hall 22, Meier 26, Whitmont/Perera 28); sec03/04 gaps re-verified (0 rescues); AGENTS.md: EN sweep procedure documented
This commit is contained in:
parent
6334148b45
commit
1ff76b4270
23 changed files with 114 additions and 23 deletions
38
AGENTS.md
38
AGENTS.md
|
|
@ -439,6 +439,44 @@ Special cases:
|
|||
- ~4s between libgen requests, ~3s between RSL pages, ~2s between shop requests.
|
||||
- All tools cache nothing by default; reruns are fine but avoid duplicate queries in a batch.
|
||||
|
||||
## EN original sweep (soft libgen queries) — validated 2026-09-20 (sec01: +50 files)
|
||||
|
||||
Exact full-title libgen queries miss a lot (indexed titles vary: reprints, e-books, variants).
|
||||
The validated approach for closing the EN download gap of a section:
|
||||
|
||||
1. **Build the gap list** from `downloads/0N-*/MANIFEST.md` («EN full text unavailable /
|
||||
not present» block) + SUMMARY Table 3 DL EN column (rows with «— (restricted)» / «—»).
|
||||
2. **Soft libgen queries, 2–3 rounds** (author-centric, not title-guessing):
|
||||
- round 1: `<author surname> <1–2 distinctive title words>`;
|
||||
- round 2 (0-hit items): `<author surname>` alone, filter results by title/first name;
|
||||
- round 3: alternate wording (shortened title, co-author, publisher word).
|
||||
Whole-word matching still applies (full words, case-insensitive). `tools/lg.py` prints
|
||||
`(unparsed…)` on 0-result pages — treat as 0 but confirm once with raw curl grep `edition.php?id=`.
|
||||
3. **Collect candidates → batch metadata:** `https://libgen.vg/json.php?object=e&addkeys=*&ids=<comma-list>`
|
||||
(≤10 ids) → `files` dict gives `f_id`+`md5` per file; then `object=f&addkeys=*&ids=<f_ids>`
|
||||
for extension/filesize. Filter by year/edition to match the reading-list book (identity check:
|
||||
title + author + era; e.g. SE15 vs SE22 look similar in truncated titles).
|
||||
4. **Download queue (serial):** `tools/lgdl.py dl <f_id> <outdir> <name>`; resume works via `.part`
|
||||
(curl `-C -`); the tool retries with fresh get.php keys. Huge files (>100 MB) can crawl at
|
||||
~10 KB/s on libgen CDN (random stuck offsets) — let it stall-detect, or abandon in favor of
|
||||
another edition of the same book (e.g. Great Mother: 643 MB Princeton 2015 abandoned →
|
||||
64 MB Bollingen 1991).
|
||||
5. **Monitor-kill trap (2026-09-20):** NEVER `pkill -f "lgdl.py"` (or any pattern that appears in
|
||||
the queue bash's own command line) — it kills the queue itself. Kill only the worker by comm:
|
||||
`ps -eo pid,comm,cmd | awk '$2=="python3" && /tools\/lgdl/ {print $1}' | xargs -r kill`.
|
||||
6. **Verify after:** magic bytes (`%PDF` / `PK\x03\x04` for epub; discard HTML stubs) +
|
||||
`pdfinfo` page count for small files — 2–14 pp «pdf» are fragments/essays, keep but mark
|
||||
«фрагмент» in MANIFEST (e.g. Way of All Women 1933 = 2-pp, Self in Transformation = 14-pp JAP essay).
|
||||
7. **archive.org (ia) as complement** for trade-restricted items: `tools/ia.py search "…"` →
|
||||
open items only (`RESTRICTED` = borrow-only). Jung CW: single open item `CarlJungCollectedWorks`
|
||||
(vols 1–18 + seminars, PDF+EPUB, direct `/download/<item>/<file>`).
|
||||
8. **Docs to update after:** SUMMARY Table 3 DL EN + totals row, MANIFEST (EN section with lg f_id /
|
||||
ia item per file; «not downloadable» list), per-card line «EN downloaded (date): <source> — <edition>»
|
||||
(or «EN full text: NOT FOUND (date, 3 rounds)» for the true gaps), CONTEXT.md note. One commit per section.
|
||||
|
||||
Result sec01: 50 files from 58 missing (8 genuinely absent: 13, 16, 38, 42, 48, 50, 52, 54 +
|
||||
643 MB 24a replaced by 1991 twin).
|
||||
|
||||
## Download phase conventions (user rule, 2026-07-18)
|
||||
|
||||
- **EN originals: libgen.vg FIRST** (the EN collection is bigger than the RU one — most trade
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue