Most popular now

Why AI Plagiarism Is So Hard to Prove, According to MIT

Illustration of plagiarism research in IT
Використання штучного інтелекту при створенні контенту ускладнює виявлення плагіату, як пояснюють експерти з MIT. Photo: НВ — Техно

Generative AI and the Attribution Puzzle

According to НВ — Техно: New findings from MIT CSAIL reveal that the more data a generative AI model is trained on, the harder it becomes to identify which specific images influenced a particular output. Study lead Zheng Dai says that when a model generates an image, there is a natural impulse to point to the exact portion of training data behind it-but this new work shows that doing so grows more difficult as the dataset expands.

The team built about two dozen diffusion models of their own, training them on datasets ranging from a few hundred images to hundreds of thousands. They evaluated the models using facial photos and works from various artists. Across all tests, larger training sets made it harder to determine the contribution of any single piece. The researchers label this effect "attribution decay."

Method, Findings, and Broader Impact

Their approach compared a model's normal output with a version produced after specific images were removed from training data. The results showed that with large datasets, the link between a particular artwork and the generated result weakens. The study concludes that for large diffusion models, no reliable method exists for establishing which specific work gave rise to a certain element in an AI-generated image.

This makes it harder for artists to prove that their work directly caused an AI system to produce something. Since commercial models are far larger, the difficulty is even greater there. The authors emphasize that understanding these mechanisms is vital for continued technical research and for developing policies around AI usage.

These outcomes have substantial implications for art and copyright. They complicate efforts to determine whether or to what extent an artist's work influenced a generative model's output. With AI tools becoming increasingly common in creative industries, the push for clearer standards and fair-use regulations is likely to intensify. The legal and ethical uncertainty created by attribution decay will therefore demand careful attention.

As the challenges of attributing artistic influence in AI-generated content become increasingly complex, it's essential to explore innovative solutions. For instance, Anthropic's new approach to marking text with an invisible code provides a potential method for identifying the origins of AI outputs. This development may offer insights into how to navigate the intricate landscape of copyright and creativity in the age of generative AI.

Read also

Advertisement