遇见数据集

Disclosures-SSRC/Detecting-Access-Violations-in-a-LLMs-Pre-Training-Data

收藏
Hugging Face2025-11-18 更新2025-12-20 收录
官方服务:

资源简介:

# Beyond Public Access in LLM Pre-Training Data The official HuggingFace repository for the paper "Beyond Public Access in LLM Pre-Training Data" by [The AI Disclosures Project](https://www.ssrc.org/programs/ai-disclosures-project/). Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate whether OpenAI's large language models were trained on copyrighted content without consent.

提供机构:
Disclosures-SSRC
二维码
社区交流群
二维码
科研交流群
商业服务