Frontier AI models, like airplanes, should be required to go through technical testing and auditing, and their release should be blocked or reversed as a threat to public safety if they do not meet high standards of safety.
Dario Amodei: On the Record
- 3,907Claims
- 108Sources
- 2016-2026Window
- 91%Audio
BUY INDIVIDUAL · $149 BUY COMMERCIAL · $990 FREE SAMPLE PACK · 25 CLAIMS · ALL 6 FILES
EXHIBITS
I think there's a 25 % chance that things go really, really badly
That's right. 25 % is too high. We're trying to make that probability much, much lower. That is the goal.
Some of the early companies that we gave this to said things like, this is a super weapon. You should have to own a gun license to use it. Please don't release this.
we've been big on export controls on chips going to China because I'm very worried about what an authoritarian country would do with that kind of power.
HOW THIS WAS MADE
Every quote in the exhibits is verbatim from the source audio (or text), transcribed and speaker-verified. Timestamps index the source video. Recut and re-posted sources are cited to the original. Nothing in the record is spliced, paraphrased, or taken from an interviewer's mouth. The full method, including the verification gates every release must pass, is on the methodology page.
THE CORPUS
108 sources spanning August 2016 to August 2026. 92 are recorded appearances, from podcasts and summit panels to Senate testimony and documentaries, accounting for 3,559 claims. The other 16 are his own writing, essays, op-eds, and official statements, accounting for 348.
The cadence is accelerating: 5 sources in the seven years through 2022, then 9 in 2023, 15 in 2024, 28 in 2025, and 51 in the first eight months of 2026.
The paid package is one versioned folder: the data files, data dictionary, datasheet, errata, codebook, label-quality disclosure, audit report, license, a checksums file covering every member, and a browsable record book with per-source citation pages.
CHECKSUMS · SHA-256 OF THE DELIVERED DATA FILES
- CLAIMS CSV
- 1e8b5b66842daa18d29ba88d9c57eb47c626fc4d9fb3036ffe298446873df9ca
- CLAIMS JSONL
- 53f0dd689c8cae114d29b2f6a7b876e64c0a478a529cf9fea333c53b1e6220cb
- SOURCES CSV
- ab8e2c9957388a95b4ae800a236a6ae3c1a9089fb96d6428e676590aef8010ef
- SOURCES JSONL
- 51e0f2b576d2ae615f5e265d4f8dfea168a52c7b65d3893a36d1ed4ce0e0989d
- ENTITIES CSV
- d01a9606910132e937449bcdcf04de8e3b60a77fc54217414dc090b1e677fa2b
- ENTITIES JSONL
- e427437f2c410f88bf8f95e54b20c4de38e31573d05d792ef577b31eec6b85d2
FIELD GUIDE · WHAT EACH COLUMN MEANS
- date
- appearance or publication date (YYYY-MM-DD)
- source
- the specific appearance or writing ("Lex Fridman Podcast #452", "The Adolescence of Technology"); every claim from it shares this value. Matches channel when a show is self-titled
- channel
- the outlet or person who published it ("Lex Fridman", "Bloomberg Originals", "CBS News")
- source_type
- audio (spoken: interview, talk, testimony) or text (written: essay, op-ed, statement)
- source_format
- genre of the source: essay, op-ed, statement, or testimony (text); podcast, interview, panel, keynote, fireside, or documentary (audio). interview = broadcast or press one-on-one; fireside = on-stage moderated conversation at an event
- timestamp
- position in the recording; blank for text sources, which have no timecode. Two-part values are minutes:seconds with minutes unbounded (67:26 = 67 minutes); three-part values are H:MM:SS
- proposition
- the claim as a standalone statement of his actual position; reads correctly on its own
- quote
- verbatim excerpt the claim is grounded in
- claim_type
- factual, evaluative, normative, causal, definitional, reportative, comparative, or untyped
- polarity
- affirm or deny, tagging the speech act; the proposition is already standalone-true
- temporal
- past, present, or future (future = a forward prediction); blank when undetermined
- confidence
- extraction confidence; 0.7 to 1.0 in this file (lower-confidence rows are held out for review)
- ts_source
- how the timestamp was established: verified = checked against the transcript at extraction; audio_word = re-derived from word-level audio timing; text = written source. No estimated timestamps ship
- attribution
- speaker verification: confirmed = the verbatim quote sits in a speaker-labeled segment of the recording; confirmed-text = a written source. Unverified rows are held out of this file
- speaker
- the speaker as the source identifies them. Most rows are Dario Amodei; interviewers, co-panelists, and colleagues carry their own names; co-authored writings name all authors. Empty where the record does not identify the voice: attribution is never guessed
- claim_id / source_id
- stable IDs. claim_id is unique per row; source_id is the per-source slug (date + channel + title), one per appearance: the join and audit key
- quality_flags
- mostly empty; echo = the proposition repeats the quote verbatim rather than distilling it; truncated = the quote ends mid-sentence; context_dependent = the proposition opens on an unresolved reference and needs its quote and timestamp for full context. Filter to empty for the cleanest subset
- topic_l1
- one of 14 level-1 topics, assigned against the codebook shipped in the package; the label-quality disclosure reports the measured error rate
- topic_l2
- a short emergent subtopic phrase; free text, not drawn from a fixed list
- content_flags
- semicolon-joined descriptive labels (numeric, self_biographical, commitment, characterization); they describe what kind of statement a row is, never whether it is true
- location
- canonical URL for text sources; blank for audio rows
CHANGELOG
- v2026.09.1
- 2026-09-05. 3,907 claims from 108 sources. Adds the 2026-08-26 CNBC Closing Bell Overtime sitting with Salesforce.
- v2026.08.6
- 2026-08-11. The package now includes the rendered views: a browsable record book (the full record organized by topic, readable in a browser), coverage map, computed findings, position timelines, and a per-source citation page for every source. The audit summary now explains each quality flag beside its count. No data changes; every table is byte-identical to v2026.08.5.
- v2026.08.5
- 2026-08-11. No data changes: every data file is byte-identical to v2026.08.4. The delivered audit report now states the record's verification tier as machine-verified; analyst-audited review is offered as a commissioned audit rather than an instant download. Audit re-run at gate 1.17.0.
- v2026.08.4
- 2026-08-06. 3,883 claims from 107 sources (up from 2,493 and 64), 649 entities. Topic labels and content flags on every claim, with the codebook they were assigned against and a measured label-quality disclosure shipped in the package. Speaker labels corrected against the record and entity spelling variants merged, with the alias tables included. The package now ships as one versioned folder with a computed data dictionary, datasheet, errata, audit report, and a checksums file; every document regenerates from the shipped rows and is byte-verified before release.
- v2026.08.1
- 2026-07-05. First public release: 2,493 claims from 64 sources (2,166 from recordings, 327 from his writing), 498 entities; every table in both CSV and JSONL; full-population audit passed (schema, cross-file counts, and verbatim grounding of each quote against its source).
COMMISSIONS
Need a record like this on a different subject, or on media you hold? Custom datasets, dossiers, consultation, or partnership: zac@librarianlabs.com or @zacforristall.