sec06: key-vs-f_id discovery — 12 EN files recovered (6 rescues) + precheck v2

- USER FINDING: libgen edition files map {key: {f_id, md5}} — KEY and f_id VALUE
  are DIFFERENT file records; we were downloading/prechecking by the KEY (wrong file).
- Recovered by f_id: 09-t1 (Eliade UCP 1978), 13 Sungod (Cornell 2010), 14 Solomon
  (Karnac 2007), 17 God-Image (Inner Light 1992), 19 Hornung (UNC 1982), 20 Psyche in
  Scripture (1995), 23 Dionysos (Bollingen LXV/2 Princeton) + backups 12/16/24/25/26/27.
  All content-verified. '24 systemic mislinks' = our bug, not libgen.
- lg.py: objects[]=f,e,s,a,p,w + topics[] (file-level search; found Solomon md5) +
  md5 in output. lgdl.py: bymd5 subcommand.
- precheck v2: rev-value match = OK; locator words >=2 shared = OK; words present
  0 shared = BAD; hash/ISBN/empty locator = no signal (SUSPECT); digit-runs masked
  before tokenization (hex 'aaae' trap).
- MANIFEST/SUMMARY/cards/AGENTS.md updated. 34 files, 608 MB (28 backup still downloading).
This commit is contained in:
Dmitry Kokorin 2026-09-25 10:04:01 +03:00
parent 96c0e1190d
commit b7ddfac17e
30 changed files with 298 additions and 137 deletions

View file

@ -33,10 +33,14 @@ def fetch(url, tries=3):
def main():
q = sys.argv[1]
limit = int(sys.argv[2]) if len(sys.argv) > 2 else 10
url = "https://libgen.vg/index.php?" + urllib.parse.urlencode({
"req": q, "res": "100", "dlt": "0", "ln": "0",
"columns[]": ["t","a","s","y","p","i","l","x","sz"]
})
# objects[] MUST include 'f' (files): editions-only search misses files that
# exist in libgen but are not linked to the right edition (verified 2026-09-25:
# Solomon "Self in Transformation" found by file search, missing in edition search).
params = {"req": q, "res": "100"}
params["columns[]"] = ["t","a","s","y","p","i"]
params["objects[]"] = ["f","e","s","a","p","w"]
params["topics[]"] = ["l","c","f","a","m","r","s"]
url = "https://libgen.vg/index.php?" + urllib.parse.urlencode(params, doseq=True)
html = fetch(url)
# pagination hint
m = re.search(r'page=(\d+)', html)
@ -52,8 +56,11 @@ def main():
if len(tds) < 6:
continue
# first cell is usually a link with id, title is the big cell
title = max(tds, key=len) if tds else ""
line = " | ".join(tds[:9])
# file rows: surface the md5 (ads.php link) — feed to tools/lgdl.py bymd5
md5s = re.findall(r'ads\.php\?md5=([0-9a-f]{32})', r)
if md5s:
line += f" [md5={md5s[0]}]"
print(line[:300])
count += 1
if count >= limit: