content audit sec01-06: fixes (audit_content.py + 12 file/card corrections)

- tools/audit_content.py: per-file vs item-title triage (SUSPECT/SMALL/NO-TEXT/NO-CYR)
- sec01: .part removed; 10 files restored .pdf ext (LFS rename); #04/#19/#27 verify notes
- sec02 #16: corrupted double-encoded txt -> clean flib fb2 b/844509; #17/#22 doc verified via catdoc
- sec03 #11: -b epub = Spanish von Franz (Paidós 1983, forged EN OPF) -> bonus label
- sec03 #42: КАРО 2012 'Irish Tales' = EN reader (not RU, not the listed book) -> bonus rename
- sec03 #56: Onians b/c/d/e = Cambridge 24-87KB previews -> labeled
- sec06 #18: filename 1981->1989 (Perera 'Descent to the Goddess' 1989; 1981 = other book)
- sec06 #28: ia Zimmer OCR layer = foreign Devanagari text; images = RKP (p.100 verified)
- sec04 #30: buksmart 2020 = RU/FR bilingual w/ VeryPDF watermarks -> note
- linter: azw3/mobi ext + underscore in fname regex; manifest cross-check ignores off-disk rows;
  8 legacy cards section-order fixed (CR before DL)
- AGENTS.md: 'Content audit (2026-09-25)' lessons section
- check-md: 0/219
This commit is contained in:
Dmitry Kokorin 2026-09-26 01:32:02 +03:00
parent 85c41acd3c
commit f5cb34195e
44 changed files with 20976 additions and 83 deletions

View file

@ -19,7 +19,7 @@ def parse_manifest_files(manifest_path):
if not os.path.exists(manifest_path):
return items
c = open(manifest_path, encoding='utf-8').read()
fname_re = r'([0-9]{2}-[a-zA-Z0-9\-]+(?:\.[a-zA-Z0-9]+)*\.(?:pdf|fb2|epub|doc|docx|chm|djvu|txt|zip))'
fname_re = r'([0-9]{2}-[a-zA-Z0-9_\-]+(?:\.[a-zA-Z0-9]+)*\.(?:pdf|fb2|epub|doc|docx|chm|djvu|txt|zip|mobi|azw3))'
# sec01/02: table rows | file | size | source | note |
for m in re.finditer(r'^\|\s*' + fname_re + r'\s*\|\s*([^|]+)\|\s*([^|]+)\|', c, re.M):
f, size, src = m.group(1), m.group(2).strip(), m.group(3).strip()