Epstein files: How victims remain exposed to identification
US officials vowed to revise redactions after files exposed information that identified survivors of sexual abuse by Jeffrey Epstein and his associates. A DW investigation
US officials vowed to revise redactions after files exposed information that identified survivors of sexual abuse by Jeffrey Epstein and his associates. A DW investigation found that sensitive information remains online. "I cannot live without looking over my shoulder," a sexual abuse survivor told US lawmakers in May, describing how her personal information was publicly exposed in January, when the Department of Justice (DOJ) released hundreds of thousands of documents related to the financier Jeffrey Epstein. "I can only imagine the long-term impact this mistake will have on my life." Three days after the DOJ released the largest of 12 datasets on Epstein on January 30, a department lawyer told a federal court that "human or technical error" had been among the factors that led to the publishing of personally identifying information of people who had been sexually abused by Epstein and his associates. As a result, the DOJ said it would remove, review and redact flagged documents before republishing them. DW's investigative unit, along with its Innovation and Data teams, spent months reviewing thousands of documents released by the DOJ. In the six months since the release, DW identified dozens of files that still contained information — including names, faces and email addresses — that could be used to identify survivors, witnesses and informants, among other people. DOJ Epstein Library By mid-February, DW had scraped more than 800,000 files from the DOJ's Epstein Library — a challenging task with the archive constantly changing: New files were being added, others were taken down and modified files were popping up after days of absence. But it wasn't until DW obtained an earlier set of files archived by a publicly facing data preservation project on GitHub that we realized more than 500,000 files were removed between the initial days of the release and our scrape.
To verify the contents of the original release, DW generated a unique digital fingerprint, known as a hash, for every document. Hashes confirm that a document matches the original and hasn't been altered. By comparing these fingerprints with documents scraped from the DOJ website about two weeks later, DW could determine which files matched, as well as identify documents that had been removed from the DOJ website. DW also found that some files had changed because the digital fingerprint was no longer the same: They had the same file name but were larger or smaller than the original. When DW reviewed this subset, we found that they represented files that had been further redacted — or information was added. In one instance, a survivor described how Epstein had abused her for years as a minor in her testimony to lawyers. Her name was redacted throughout the transcript until the last page, when a lawyer thanked her for her testimony and said her name, and it remains unredacted as of publication. Her exposed name was not an isolated case. US lawmakers created strict rules on how the Epstein files should be published, such as ensuring that personally identifying information of victims be redacted to protect them from public exposure Image: picture-alliance/dpa/M. Reynolds Deep audio analysis To analyze audio files, including victim statements and tipoffs, DW's investigative unit collaborated with the Fraunhofer Institute for Digital Media Technology's Media Distribution and Security research group in Ilmenau, Germany. The MDS specializes in audio authenticity, manipulation and speech synthesis detection. After removing duplicated files from the publicly facing data preservation project and DW's scrape, we were left with more than 150 audio files, which the MDS's forensic audio researchers analyzed using state-of-the-art tools developed at the institute.
