News thumbnail
Technology / Fri, 28 Aug 2026 The Hindu

How massive data pools remove the “smoking gun” of AI training

In small datasets, an artist’s work could be a significant addition. They can now argue that even if they had never seen a specific artist’s work, their model would have produced a nearly identical result anyway. While an autoregressive model predicts the next “token” in a sequence, a diffusion model is learning the geometric and semantic spread of visual concepts. For instance, a lab could train its diffusion model from synthetic data generated from outputs created by the current crop of AI models. AI generated images and videos from these newer models will contain far fewer data that can be sourced back to specific artists.

Even as AI firms expand model capabilities, frontier labs continue to navigate a legal reckoning sparked by the first wave of copyright lawsuits that threatened to derail the generative AI boom.

The argument from artists has been that AI output that looks like a specific piece of art is because that work was “stolen” during training the model. But a landmark research paper, ‘Outputs of Generative Diffusion Models are Often Unattributable,’ by MIT Computer Science and Artificial Intelligence Laboratory (CSAIL) is effectively dismantling the scientific basis for that claim, suggesting that in the world of frontier models, visual similarity is not the same thing as causal theft.

The paper’s findings could make the legal teams of artists uneasy. By using “ablatable ensembles” — a machine learning method than can systematically remove specific training data to test outputs — researchers have found that as training datasets grow, the influence of any single image or artist drops toward zero.

In small datasets, an artist’s work could be a significant addition. Removing it would collapse the output. But in a large-scale model, that same artist’s work would be just a drop in an ocean. It is quite likely that the model could have learned the concept of “impressionism” or “cyberpunk” from so many derived sources that the removal of any one data point leaves the output virtually unchanged.

A shield for diffusion models

The paper showed that an AI’s reliance on specific training samples decays predictably as the data pool expands. For AI labs like Midjourney or OpenAI, this research provides a powerful shield. They can now argue that even if they had never seen a specific artist’s work, their model would have produced a nearly identical result anyway.

The research focuses specifically on Diffusion Models — the tech behind DALL-E, Stable Diffusion, and Midjourney — which generate images by iteratively removing noise from a canvas until a coherent picture emerges. These models operate differently than the autoregressive models that power Large Language Models (LLMs) like ChatGPT or Claude.

While an autoregressive model predicts the next “token” in a sequence, a diffusion model is learning the geometric and semantic spread of visual concepts. Because of this, diffusion models are less likely to memorise and store a specific file and more likely to synthesise a generalised idea.

Crucially, this specific paper does not focus on text-based outputs; the “unattributable” phenomenon is a visual one, leaving the door open for text-based copyright claims by authors where LLMs have been caught quoting books or articles verbatim.

But the diffusion-based model creators can claim legal immunity from scale. A model trained on a billion images is harder to sue than a model trained on a million as the measurable change caused by the removal of one piece of art from a large dataset is insignificant statistically.

For frontier AI labs that have built their models on diffusion-based architecture, this research will come as good news as it suggests that artistic style can get generalised as the model learns from more sources. As these models become more sophisticated, they are quite likely to get better at mimicking humans, and become increasingly independent of the individual data points that created them.

The real test

While frontier AI labs might use this research as a scientific basis to disprove the accusation from artists that their models contain “stolen art”, the real test will be in how they use it in a court of law.

If an AI generates an image that looks exactly like an artwork by one of the masters, a judge may not care whether a statistical model shows a negligible correction. But for the AI researchers building the next generation of image and video models, the message looks settled. For instance, a lab could train its diffusion model from synthetic data generated from outputs created by the current crop of AI models. AI generated images and videos from these newer models will contain far fewer data that can be sourced back to specific artists. Also, the bigger their dataset, the less they need to be worried about data theft allegations.

In the race to build advanced AI, the most valuable asset isn’t just the data itself, but the scale that makes that data anonymous.

Authors still have a “smoking gun”

While image generators might hide behind the “unattributability” of scale, authors have a much more tangible “smoking gun.” When an LLM reproduces a copyrighted paragraph word-for-word, it is far easy to match the output to a copyrighted source. In the case of LLMs, the model stores a specific sequence of words to learn how a sentence in written.

Even if a model is so large that no single book is essential to its function, authors have argued that the collective “uncompensated use” of their work is what creates the value in the first place.

Also, unlike diffusion models that generate AI content from noise, LLMs pick tokens from a fixed vocabulary that are organised under specific knowledge categories. That means it will be easier to attribute a written piece of work than an artistic style. What the AI deliver in the former case is knowledge, but in the latter style.

For now, the industry is in a holding pattern. The “unattributability” paper gives a technical edge to image generators, but for the text-based giants, the question of whether a written work can be “diluted” the way a brushstroke can remains an open question.

© All Rights Reserved.