sec01: round 2 closed — 0 new rescues; cache processed catalogs in data/catalogs/
- data/catalogs/: cogito Jungian series (518 titles, 30 pp.), litres 140409 (24 JSON-LD),
OPP/MISP reading list (55 entries) — source URLs + dates in headers
- author-level check of 16 ❌ against cogito shop search (11 author queries) + full 518-title
catalog diff: 0 matches; noise rejected (Shamdasani≠Casement, Jung 'Нераскрытая самость'
≠ McFarland-Solomon, Berdyaev essay)
- AGENTS.md: catalog caching convention
This commit is contained in:
parent
b34b80196f
commit
e7b35223af
5 changed files with 651 additions and 0 deletions
|
|
@ -257,6 +257,13 @@ sections/0N-<name>/ # one .md per BOOK (chapter-level list items are merged int
|
|||
- Anchor titles in result rows contain a raw `>` (title="Хранение<br/>Заказ") — regexes must
|
||||
not use [^>]+ across the anchor attrs.
|
||||
|
||||
## Catalog caching convention
|
||||
|
||||
Processed/scraper-unfriendly catalogs are cached under `data/catalogs/` (committed, with source
|
||||
URL + date in the header) so negative verification stays re-runnable without re-crawling.
|
||||
Examples: `cogito-yungianskaya-series-2026-07-17.md` (30 pp. CRW crawl, 518 titles),
|
||||
`litres-140409-2026-07-17.md`, `opp-misp-reading-list-2026-07-17.md`.
|
||||
|
||||
## Not usable / low value
|
||||
- libgen biblioservice worldcat/googlebooks/isbndb/udc (broken on this mirror)
|
||||
- imaton.com (МААП educational publisher, not translations)
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue