content audit sec01-06: fixes (audit_content.py + 12 file/card corrections)
- tools/audit_content.py: per-file vs item-title triage (SUSPECT/SMALL/NO-TEXT/NO-CYR) - sec01: .part removed; 10 files restored .pdf ext (LFS rename); #04/#19/#27 verify notes - sec02 #16: corrupted double-encoded txt -> clean flib fb2 b/844509; #17/#22 doc verified via catdoc - sec03 #11: -b epub = Spanish von Franz (Paidós 1983, forged EN OPF) -> bonus label - sec03 #42: КАРО 2012 'Irish Tales' = EN reader (not RU, not the listed book) -> bonus rename - sec03 #56: Onians b/c/d/e = Cambridge 24-87KB previews -> labeled - sec06 #18: filename 1981->1989 (Perera 'Descent to the Goddess' 1989; 1981 = other book) - sec06 #28: ia Zimmer OCR layer = foreign Devanagari text; images = RKP (p.100 verified) - sec04 #30: buksmart 2020 = RU/FR bilingual w/ VeryPDF watermarks -> note - linter: azw3/mobi ext + underscore in fname regex; manifest cross-check ignores off-disk rows; 8 legacy cards section-order fixed (CR before DL) - AGENTS.md: 'Content audit (2026-09-25)' lessons section - check-md: 0/219
This commit is contained in:
parent
85c41acd3c
commit
f5cb34195e
44 changed files with 20976 additions and 83 deletions
24
AGENTS.md
24
AGENTS.md
|
|
@ -531,6 +531,30 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
|
|||
queue bash (dual-queue downloads corrupt the cookie jar). AA (annas-archive.gl) = 403 bot-wall
|
||||
for anon downloads from this box; i2p not worth the router setup.
|
||||
|
||||
## Content audit (2026-09-25) — what the full-file vs item-title sweep found
|
||||
|
||||
Tool: `tools/audit_content.py <sec>...` (samples first/mid pages of every file, scores title
|
||||
overlap, flags SUSPECT/SMALL/NO-TEXT/NO-CYR). Ran on all sections 01–06 (sec05: fix the
|
||||
per-item header parse first — sec05 uses «— ✅» not «— N files»). Findings (commits
|
||||
`85c41ac` + follow-ups):
|
||||
1. **Mislinked whole files caught:** sec01 #29 RU fb2 = fantasy novel; sec01 #47 two review
|
||||
clippings (1950 «About Books»); sec03 #11 `-b` epub = Spanish von Franz (Paidós 1983) with
|
||||
forged EN OPF title; sec03 #42 «КАРО 2012» = English reader (Irish Tales), not the listed
|
||||
book and not a RU translation; sec02 #16 bosnak .txt = double-encoded mojibake (replaced by
|
||||
clean flib fb2 b/844509).
|
||||
2. **Foreign OCR layer (file OK, text layer wrong):** sec06 #28 ia Zimmer = all-Devanagari OCR
|
||||
(Lal Bahadur Shastri Academy stamp); images verified = RKP «Philosophies of India» (p.100
|
||||
«PART II Philosophies of Time»). Keep; the twin lg file is the clean copy.
|
||||
3. **Previews/fragments labeled as editions:** sec03 #56 Onians b/c/d/e = Cambridge 24–87 KB
|
||||
chapter previews (marked ПРЕВЬЮ in MANIFEST); sec01 #12 was a 23-pp PagePlace preview (→ full).
|
||||
4. **False positives to expect:** windows-1251/UTF-16/KOI8 fb2/txt (decoding, not corruption),
|
||||
image scans without OCR (NO-TEXT/NO-CYR — verify by rendering pages: pdftoppm/ddjvu +
|
||||
imagemagick), spaced-out OCR titles («T H C T F»), library stamps (Cornell/India/Claremont),
|
||||
VeryPDF watermark pages (sec04 #30 buksmart — watermark + FR part, legit 2020 ed.).
|
||||
5. **Bonus rule in practice:** different-but-related books stay as bonus files with
|
||||
`Бонус: ДРУГАЯ книга (...)` note, status unchanged; rename file to drop the false lang/year
|
||||
label. Same for wrong-year filename (sec06 #18 «1981» → 1989, Perera had two «Descent…» books).
|
||||
|
||||
## Catalog caching convention
|
||||
|
||||
Processed/scraper-unfriendly catalogs are cached under `data/catalogs/` (committed, with source
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue