sec06: key-vs-f_id discovery — 12 EN files recovered (6 rescues) + precheck v2

- USER FINDING: libgen edition files map {key: {f_id, md5}} — KEY and f_id VALUE
  are DIFFERENT file records; we were downloading/prechecking by the KEY (wrong file).
- Recovered by f_id: 09-t1 (Eliade UCP 1978), 13 Sungod (Cornell 2010), 14 Solomon
  (Karnac 2007), 17 God-Image (Inner Light 1992), 19 Hornung (UNC 1982), 20 Psyche in
  Scripture (1995), 23 Dionysos (Bollingen LXV/2 Princeton) + backups 12/16/24/25/26/27.
  All content-verified. '24 systemic mislinks' = our bug, not libgen.
- lg.py: objects[]=f,e,s,a,p,w + topics[] (file-level search; found Solomon md5) +
  md5 in output. lgdl.py: bymd5 subcommand.
- precheck v2: rev-value match = OK; locator words >=2 shared = OK; words present
  0 shared = BAD; hash/ISBN/empty locator = no signal (SUSPECT); digit-runs masked
  before tokenization (hex 'aaae' trap).
- MANIFEST/SUMMARY/cards/AGENTS.md updated. 34 files, 608 MB (28 backup still downloading).
This commit is contained in:
Dmitry Kokorin 2026-09-25 10:04:01 +03:00
parent 96c0e1190d
commit b7ddfac17e
30 changed files with 298 additions and 137 deletions

View file

@ -11,10 +11,10 @@ We work **section by section**. Sections **01–05 are COMPLETE** (01: 38✅/2
2026-09-22 (7 rescues from user tips — see Lessons), **EN downloads COMPLETE 2026-09-23**
(30 files, 362 MB; 5 EN gaps unavailable + 2 ia-restricted; SUMMARY.md at-a-glance table done).
Section 06: research + user oracle DONE 2026-09-24 (28 карточек: **19 ✅ / 1 🔶 / 8 ❌**;
7 rescues; каталог Касталии 1285 в data/catalogs + rule в gate v3); **downloads COMPLETE
2026-09-25** (22 файла, 413 MB: RU 13 + EN 8 + DE 1 бонус; 6 кросс-рефов на sec01-03;
RU платные castalia-PDF не скачивались; EN: libgen-батч почти весь mislinked → 9 файлов с ia
open-uploads; sec03#20 RU-дыра пометка в MANIFEST). Section 07 pending; 08–12 scope
7 rescues; каталог Касталии 1285 в data/catalogs + rule в gate v3); **downloads FINAL
2026-09-25** (34 файла, 608 MB: RU 12 + EN 21 + DE 1 бонус; 6 кросс-рефов на sec01-03;
EN: 9 с ia open-uploads + 12 с libgen после находки key-vs-f_id — rescues 09-t1, 13, 14, 17, 19, 20, 23;
RU платные castalia-PDF не скачивались; sec03#20 RU-дыра пометка в MANIFEST). Section 07 pending; 08–12 scope
decision pending.
## Repository layout
@ -167,6 +167,11 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
→ Use **full inflected words** (nominative, as stored in titles) and 2–3 distinctive words.
- Author search works (`req=Нойманн`) but pulls homonyms — filter by first name in results.
- Pagination: `&page=N` works.
- **Search objects (2026-09-25):** `objects[]` param = что искать (f=files, e=editions, s=series…).
Без него = editions-only — файлы, не привязанные к нужному изданию, НЕ находятся
(Solomon: editions-поиск = только mislinked 2021-издание, files-поиск = настоящий файл 2018-го).
lg.py теперь шлёт `objects[]=f,e,s,a,p,w` + `topics[]=l,c,f,a,m,r,s` и печатает `[md5=…]` в строках
→ `tools/lgdl.py bymd5 <md5>` = f_id + метаданные. Это нашло все 6 rescues sec06.
- Edition page: `https://libgen.vg/edition.php?id=<id>`.
### 3. cogito-shop.com — current Russian trade (Jungian/psychoanalysis specialists)
@ -473,30 +478,40 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
4. **flib authorall for prolific RU authors** (Элиаде a/2943 = 49 books) caught «Тайные общества»
(sec05 #13) that libgen title search missed. Always run flib author sweep for authors with
>5 RU books.
5. **libgen mislinked-file trap (SYSTEMIC — 9/20 files in the sec05 RU batch were wrong)**:
edition files can be mislinked (141685581 «На пороге инициации» carried an ePubLibre
Spanish Lermontov file; the sec05 «pdf» batch carried English academic papers:
*Food Engineering Research*, *Surface & Coatings Technology*, *International Journal of Food
Sciences*, an EN novel, even a RU fantasy «Подмененный» as fb2 for «Священное и мирское»;
the 4 van Gennep epubs were Spanish fiction). Sizes pass lgdl's md5 check. **Mandatory:
after every libgen download, content-verify: `pdfinfo` page count vs RSL record + first-page
text (pdftotext -l 3 / fb2 body head / epub xhtml head) must contain the RU title/author.**
5. **libgen mislinked-file trap + KEY vs f_id (2026-09-25, уточнение):**
(a) Edition files can be mislinked (141685581 «На пороге инициации» carried an ePubLibre
Spanish Lermontov file; the sec05 «pdf» batch carried English academic papers; the 4 van Gennep
epubs were Spanish fiction). Sizes pass lgdl's md5 check. **Mandatory: after every libgen
download, content-verify: `pdfinfo` page count vs RSL record + first-page text
(pdftotext -l 3 / fb2 body head / epub xhtml head) must contain the RU/EN title/author.**
For recent books prefer flib (clean files); flib `/b/<id>/download` → 302 → direct
`static.flibusta.is` URL.
(b) **KEY vs f_id (системная ловушка, наш баг, не libgen):** в JSON-карточке издания
`files` = `{public_key: {f_id: X, md5: Y}}` — КЕЙ и f_id-ЗНАЧЕНИЕ = РАЗНЫЕ файловые записи.
`object=f&ids=<key>` возвращает устаревший/чужой файл (Solomon: key 92709688 = 20MB рус. антология;
f_id 93253044 = настоящая 2MB-книга). Утром 2026-09-25 «24 mislinked EN» = качали по key;
скачав по f_id-значениям — rescues 06-й секции: 09-t1, 13, 14, 17, 19, 20, 23 + бэкапы.
**Правило: качать и пречекать ТОЛЬКО по f_id-значению из files-мапы (md5 для контроля).**
Файл-уровневый поиск (`objects[]=f` в lg.py) находит файлы, не привязанные к «нашему» изданию
(так найден md5 Solomon, когда editions-поиск его не показывал).
6. **Collection-inclusion questions** (#27/#28): download the smallest fb2, parse the TOC
(`<section name="title">`), grep for the EN book title + RU calque. «Символ и ритуал» (Наука
1983) CONTAINS «Ритуальный процесс» (full text, core section) but NOT «Лес символов»
(citations only).
7. **flib pdf download**: `/b/<id>/pdf` may return an HTML stub; `/b/<id>/download` → 302 →
direct `static.flibusta.is:443/b.usr/<File>.<ext>` URL — curl that directly.
8. **Content-verify gate (2026-09-22, after the sec05 mislinked batch)**: every downloaded file
gets a two-second check — magic bytes + `pdfinfo` pages + first-page/head text containing the
expected RU title/author. A file that fails is discarded and re-sourced (usually flib).
**Pre-download gate: `tools/lgdl.py precheck <f_id> <edition_id> [pages]`** (exit 0 OK /
1 SUSPECT / 2 BAD) — checks the file record's `editions` reverse-map (must contain the target
edition), `locator` batch path, size sanity; `dl <f_id> <dir> <name> <edition_id>` runs it
automatically and aborts on BAD. Mislink = skip that f_id → next file of the same edition →
next candidate edition → flib; log rejected f_ids in MANIFEST notes.
8. **Content-verify gate (2026-09-22, после sec05 mislinked batch; precheck дообучен 2026-09-25):**
every downloaded file gets a two-second check — magic bytes + `pdfinfo` pages + first-page/head
text containing the expected RU title/author. A file that fails is discarded and re-sourced
(usually flib). **Pre-download gate: `tools/lgdl.py precheck <f_id> <edition_id> [pages]`**
(exit 0 OK / 1 SUSPECT / 2 BAD). Логика (2026-09-25): (1) reverse-map file-записи: значения
= PUBLIC edition ids — совпадение с целью = сильный OK-сигнал (для записи настоящего f_id);
(2) locator: ≥2 общих слова с целью = OK; слова есть, но 0 общих = BAD (именует другую книгу);
хэш/ISBN/пустой locator = без сигнала (SUSPECT) — digit-содержащие runs маскируются до
токенизации (иначе hex даёт фейк-токены вроде «aaae»). Контент-верификация после даунлоуда —
всегда конечный гейт. Mislink = skip → next f_id of same edition → next edition → flib;
отклонённые f_id логируются в MANIFEST notes. `dl <f_id> <dir> <name> <edition_id>` запускает
precheck автоматически.
9. **EN-queue infra (2026-09-23, sec05 EN sweep)**: the libgen bottleneck is (a) `ads.php`
card-page IP-throttling under sustained load — pause 30–60 min, do NOT hammer (deepens the
block); (b) wedged CDN edges — each `get.php` 302s to a different storage host, so