content audit sec01-06: fixes (audit_content.py + 12 file/card corrections)

- tools/audit_content.py: per-file vs item-title triage (SUSPECT/SMALL/NO-TEXT/NO-CYR)
- sec01: .part removed; 10 files restored .pdf ext (LFS rename); #04/#19/#27 verify notes
- sec02 #16: corrupted double-encoded txt -> clean flib fb2 b/844509; #17/#22 doc verified via catdoc
- sec03 #11: -b epub = Spanish von Franz (Paidós 1983, forged EN OPF) -> bonus label
- sec03 #42: КАРО 2012 'Irish Tales' = EN reader (not RU, not the listed book) -> bonus rename
- sec03 #56: Onians b/c/d/e = Cambridge 24-87KB previews -> labeled
- sec06 #18: filename 1981->1989 (Perera 'Descent to the Goddess' 1989; 1981 = other book)
- sec06 #28: ia Zimmer OCR layer = foreign Devanagari text; images = RKP (p.100 verified)
- sec04 #30: buksmart 2020 = RU/FR bilingual w/ VeryPDF watermarks -> note
- linter: azw3/mobi ext + underscore in fname regex; manifest cross-check ignores off-disk rows;
  8 legacy cards section-order fixed (CR before DL)
- AGENTS.md: 'Content audit (2026-09-25)' lessons section
- check-md: 0/219
This commit is contained in:
Dmitry Kokorin 2026-09-26 01:32:02 +03:00
parent 85c41acd3c
commit f5cb34195e
44 changed files with 20976 additions and 83 deletions

View file

@ -531,6 +531,30 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
queue bash (dual-queue downloads corrupt the cookie jar). AA (annas-archive.gl) = 403 bot-wall
for anon downloads from this box; i2p not worth the router setup.
## Content audit (2026-09-25) — what the full-file vs item-title sweep found
Tool: `tools/audit_content.py <sec>...` (samples first/mid pages of every file, scores title
overlap, flags SUSPECT/SMALL/NO-TEXT/NO-CYR). Ran on all sections 01–06 (sec05: fix the
per-item header parse first — sec05 uses «— ✅» not «— N files»). Findings (commits
`85c41ac` + follow-ups):
1. **Mislinked whole files caught:** sec01 #29 RU fb2 = fantasy novel; sec01 #47 two review
clippings (1950 «About Books»); sec03 #11 `-b` epub = Spanish von Franz (Paidós 1983) with
forged EN OPF title; sec03 #42 «КАРО 2012» = English reader (Irish Tales), not the listed
book and not a RU translation; sec02 #16 bosnak .txt = double-encoded mojibake (replaced by
clean flib fb2 b/844509).
2. **Foreign OCR layer (file OK, text layer wrong):** sec06 #28 ia Zimmer = all-Devanagari OCR
(Lal Bahadur Shastri Academy stamp); images verified = RKP «Philosophies of India» (p.100
«PART II Philosophies of Time»). Keep; the twin lg file is the clean copy.
3. **Previews/fragments labeled as editions:** sec03 #56 Onians b/c/d/e = Cambridge 24–87 KB
chapter previews (marked ПРЕВЬЮ in MANIFEST); sec01 #12 was a 23-pp PagePlace preview (→ full).
4. **False positives to expect:** windows-1251/UTF-16/KOI8 fb2/txt (decoding, not corruption),
image scans without OCR (NO-TEXT/NO-CYR — verify by rendering pages: pdftoppm/ddjvu +
imagemagick), spaced-out OCR titles («T H C T F»), library stamps (Cornell/India/Claremont),
VeryPDF watermark pages (sec04 #30 buksmart — watermark + FR part, legit 2020 ed.).
5. **Bonus rule in practice:** different-but-related books stay as bonus files with
`Бонус: ДРУГАЯ книга (...)` note, status unchanged; rename file to drop the false lang/year
label. Same for wrong-year filename (sec06 #18 «1981» → 1989, Perera had two «Descent…» books).
## Catalog caching convention
Processed/scraper-unfriendly catalogs are cached under `data/catalogs/` (committed, with source