Work / gls:work:7a3f1c5d-2b89-4e6c-9d04-3f8a5b1e7c92
Declassified Document Corpus (CREST, NARA, Black Vault, NSArchive, DOE)
A bulk-available catalogue of public-domain declassified government documents — CIA CREST (incl. Stargate/remote-viewing), the NARA AWS Catalog and JFK release tranches, the National Security Archive (GWU free site), DOE OpenNet, and The Black Vault curated FOIA archive — with full provenance and per-unit rights posture. The Great Library holds the inventory; SISO Knowledge holds the durable corpus; Foundry owns source discovery; Evidence Engines perform source-grounded transformation.
Source & upstream links
Relationships
The Great Library of SISOThe Great Library of SISO holds the inventory catalog and lineage for this corpus; this Work's source-inventory record lives at registry/source-inventories/declassified-corpus-2026-08-07.json under the Great Library registry. The Great Library stays a catalog and not a warehouse: it records what exists, not the bytes.
SISO KnowledgeSISO Knowledge is the intended durable home for the corpus bytes once per-unit rights clear and the campaign advances from unverified to verified. Ingestion will follow the same shape the SISO Book Library used for Project Gutenberg: metadata index in the Great Library, payload distribution as release assets rather than mirrored into git.
SISO FoundrySISO Foundry owns source discovery for this corpus — identifying additional declassified sources, evaluating mirror candidates, surfacing rights posture changes (e.g. DOE OpenNet RD/FRD withdrawals, Black Vault archive updates), and routing primary documents into the durable corpus held by SISO Knowledge. The Foundry also evaluates whether claimable intelligence emerges from source-grounded transformation performed by Evidence Engines.
Evidence & receipts
Internet Archive CREST-1 verified: 195.5 GB, 250,000 declassified files, Public Domain Mark 1.0, TAR+CSV+torrent bundle. Combined with CREST-2/3/4 the four-torrent set covers 934,739 documents / ~12M pages / 528.7 GB.Open receipt ↗
NARA Catalog dataset on AWS verified: anonymous S3 sync with 'aws s3 sync s3://nara-national-archives-catalog/ [dest] --no-sign-request' command documented; over 261 GB of JSON metadata covering URLs for over 148 million digital objects as of 2026-03-01.Open receipt ↗
JFK Assassination Records bulk-download page verified with direct ZIPs for 2017/2018/2021/2022/2023/2025 tranches, MD5 checksums, and total corpus 'over six million pages.'Open receipt ↗
Project's own validator (npm run validate) executes against schemas/source-inventory.schema.json with internal $ref resolution, no Ajv dependency. Initial run on the declassified corpus inventory returned 3 errors: campaign owner_work_id pattern mismatch, preservation missing dirty_entries, and campaign owner Work not resolved — all three fixed by authoring this Work and adding 'dirty_entries: 0' to the inventory's preservation block.node scripts/validate.mjs
Provenance
- Registry source
- registry/works/declassified-corpus.json
- Origin
- siso
- License / redistribution
- NOASSERTION (pending)