LIBRARIAN LABS DOSSIER 001

Dario Amodei: On the Record

SUBJECT: DARIO AMODEI · WINDOW: AUG 2016 - AUG 2026 · PUBLISHED 2026-07-02 · UPDATED 2026-09-10

THE DATASET V2026.09.1 · FROM $149
  • 3,907Claims
  • 108Sources
  • 2016-2026Window
  • 91%Audio

EVERY CLAIM: PROPOSITION, VERBATIM QUOTE, DATE, SOURCE, TIMESTAMP, TOPIC LABEL, STABLE IDS · CLAIMS, SOURCES, AND ENTITIES IN CSV AND JSONL

BUY INDIVIDUAL · $149 BUY COMMERCIAL · $990 FREE SAMPLE PACK · 25 CLAIMS · ALL 6 FILES

INDIVIDUAL · $149: A PERSON. COMMERCIAL · $990: AN ORGANIZATION. EVERY LICENSE INCLUDES ALL FUTURE VERSIONS. INSTANT DOWNLOAD AT PURCHASE, PERMANENT LINK BY EMAIL. LICENSE TERMS.

EXHIBITS
Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety.

Policy on the AI Exponential, 2026-06-10. READ

I think there's a 25 % chance that things go really, really badly

Axios AI+ DC Summit, 2025-09-18. WATCH AT 17:45

That's right. 25 % is too high. We're trying to make that probability much, much lower. That is the goal.

Bloomberg Originals, The Circuit (Extended), 2026-06-17. WATCH AT 67:26

Some of the early companies that we gave this to said things like, this is a super weapon. You should have to own a gun license to use it. Please don't release this.

Bloomberg Originals, The Circuit, 2026-06-10. WATCH AT 33:41

we've been big on export controls on chips going to China because I'm very worried about what an authoritarian country would do with that kind of power.

Bloomberg Live, 2025-01-23. WATCH AT 08:40

HOW THIS WAS MADE

Every quote in the exhibits is verbatim from the source audio (or text), transcribed and speaker-verified. Timestamps index the source video. Recut and re-posted sources are cited to the original. Nothing in the record is spliced, paraphrased, or taken from an interviewer's mouth. The full method, including the verification gates every release must pass, is on the methodology page.

THE CORPUS

108 sources spanning August 2016 to August 2026. 92 are recorded appearances, from podcasts and summit panels to Senate testimony and documentaries, accounting for 3,559 claims. The other 16 are his own writing, essays, op-eds, and official statements, accounting for 348.

The cadence is accelerating: 5 sources in the seven years through 2022, then 9 in 2023, 15 in 2024, 28 in 2025, and 51 in the first eight months of 2026.

The paid package is one versioned folder: the data files, data dictionary, datasheet, errata, codebook, label-quality disclosure, audit report, license, a checksums file covering every member, and a browsable record book with per-source citation pages.

CHECKSUMS · SHA-256 OF THE DELIVERED DATA FILES
CLAIMS CSV
1e8b5b66842daa18d29ba88d9c57eb47c626fc4d9fb3036ffe298446873df9ca
CLAIMS JSONL
53f0dd689c8cae114d29b2f6a7b876e64c0a478a529cf9fea333c53b1e6220cb
SOURCES CSV
ab8e2c9957388a95b4ae800a236a6ae3c1a9089fb96d6428e676590aef8010ef
SOURCES JSONL
51e0f2b576d2ae615f5e265d4f8dfea168a52c7b65d3893a36d1ed4ce0e0989d
ENTITIES CSV
d01a9606910132e937449bcdcf04de8e3b60a77fc54217414dc090b1e677fa2b
ENTITIES JSONL
e427437f2c410f88bf8f95e54b20c4de38e31573d05d792ef577b31eec6b85d2
FIELD GUIDE · WHAT EACH COLUMN MEANS
date
appearance or publication date (YYYY-MM-DD)
source
the specific appearance or writing ("Lex Fridman Podcast #452", "The Adolescence of Technology"); every claim from it shares this value. Matches channel when a show is self-titled
channel
the outlet or person who published it ("Lex Fridman", "Bloomberg Originals", "CBS News")
source_type
audio (spoken: interview, talk, testimony) or text (written: essay, op-ed, statement)
source_format
genre of the source: essay, op-ed, statement, or testimony (text); podcast, interview, panel, keynote, fireside, or documentary (audio). interview = broadcast or press one-on-one; fireside = on-stage moderated conversation at an event
timestamp
position in the recording; blank for text sources, which have no timecode. Two-part values are minutes:seconds with minutes unbounded (67:26 = 67 minutes); three-part values are H:MM:SS
proposition
the claim as a standalone statement of his actual position; reads correctly on its own
quote
verbatim excerpt the claim is grounded in
claim_type
factual, evaluative, normative, causal, definitional, reportative, comparative, or untyped
polarity
affirm or deny, tagging the speech act; the proposition is already standalone-true
temporal
past, present, or future (future = a forward prediction); blank when undetermined
confidence
extraction confidence; 0.7 to 1.0 in this file (lower-confidence rows are held out for review)
ts_source
how the timestamp was established: verified = checked against the transcript at extraction; audio_word = re-derived from word-level audio timing; text = written source. No estimated timestamps ship
attribution
speaker verification: confirmed = the verbatim quote sits in a speaker-labeled segment of the recording; confirmed-text = a written source. Unverified rows are held out of this file
speaker
the speaker as the source identifies them. Most rows are Dario Amodei; interviewers, co-panelists, and colleagues carry their own names; co-authored writings name all authors. Empty where the record does not identify the voice: attribution is never guessed
claim_id / source_id
stable IDs. claim_id is unique per row; source_id is the per-source slug (date + channel + title), one per appearance: the join and audit key
quality_flags
mostly empty; echo = the proposition repeats the quote verbatim rather than distilling it; truncated = the quote ends mid-sentence; context_dependent = the proposition opens on an unresolved reference and needs its quote and timestamp for full context. Filter to empty for the cleanest subset
topic_l1
one of 14 level-1 topics, assigned against the codebook shipped in the package; the label-quality disclosure reports the measured error rate
topic_l2
a short emergent subtopic phrase; free text, not drawn from a fixed list
content_flags
semicolon-joined descriptive labels (numeric, self_biographical, commitment, characterization); they describe what kind of statement a row is, never whether it is true
location
canonical URL for text sources; blank for audio rows
CHANGELOG
v2026.09.1
2026-09-05. 3,907 claims from 108 sources. Adds the 2026-08-26 CNBC Closing Bell Overtime sitting with Salesforce.
v2026.08.6
2026-08-11. The package now includes the rendered views: a browsable record book (the full record organized by topic, readable in a browser), coverage map, computed findings, position timelines, and a per-source citation page for every source. The audit summary now explains each quality flag beside its count. No data changes; every table is byte-identical to v2026.08.5.
v2026.08.5
2026-08-11. No data changes: every data file is byte-identical to v2026.08.4. The delivered audit report now states the record's verification tier as machine-verified; analyst-audited review is offered as a commissioned audit rather than an instant download. Audit re-run at gate 1.17.0.
v2026.08.4
2026-08-06. 3,883 claims from 107 sources (up from 2,493 and 64), 649 entities. Topic labels and content flags on every claim, with the codebook they were assigned against and a measured label-quality disclosure shipped in the package. Speaker labels corrected against the record and entity spelling variants merged, with the alias tables included. The package now ships as one versioned folder with a computed data dictionary, datasheet, errata, audit report, and a checksums file; every document regenerates from the shipped rows and is byte-verified before release.
v2026.08.1
2026-07-05. First public release: 2,493 claims from 64 sources (2,166 from recordings, 327 from his writing), 498 entities; every table in both CSV and JSONL; full-population audit passed (schema, cross-file counts, and verbatim grounding of each quote against its source).
COMMISSIONS

Need a record like this on a different subject, or on media you hold? Custom datasets, dossiers, consultation, or partnership: zac@librarianlabs.com or @zacforristall.

QUESTIONS: ZAC@LIBRARIANLABS.COM