Can I use my Outputs to train an AI model?

DarenWatson 88 points 80 comments August 13, 2026
support.claude.com · View on Hacker News

Discussion Highlights (19 comments)

nicbou

I was not asked for permission when they trained their model on my output.

spiderfarmer

I added © 2024 - No Rights for Hypocrites to the footer of my website. Claude scraped all pages at least once every week since it came online. Now what.

ferrouswheel

You can, they are just confused about what effective altruism means.

lelanthran

Corollary: If you can't use the output to train, then you don't own them.

shikck200

They stole all the data, and then dont want you to steal it back. Its basically Robin Hood all over again.

hackmack10

Honestly, the entire dev community should save question and answers, uploaded them to a shared repo anonymously, then just use that data to distill further models and provide them to the public for free. Theoretically, this should be legal and ethical, when comparing to Anthropic's own behavior. That said, the reason you can't is Anthropic states in their terms that they don't want you to do this. Anthropic's entire business model is skating on thin ice.

ratmilk

This seems like the standard AI company hypocrisy. Hopefully Claude users disregard this nonsense.

stavros

Yeah, no. I was OK with them scraping everything if it means we get AI, but, conversely, they don't get to control what happens to their outputs. Hell, arguably they should release their weights (or at least the weights of their older models), since they trained them on the concentrated knowledge of humankind.

maaaaattttt

If more and more of the web's content is AI generated, AI companies are bound to train on each other's data. Or, what if I generate content with Claude/ChatGPT/Gemini, warp it in HTML using an open model, put this on my website conveniently dedicated to "Best practices in prompt and AI answers" for example, then train my own model that only scraps my website?

theyliesoeasily

You downloaded millions of books from shadow libraries and trained on them. I don't care about your policies, go fuck yourself.

ForgotMyUUID

> When customers use Claude to generate Outputs that then train competing models, they're essentially using our infrastructure and investment to build direct competitors to our service We did so, please do not repeat it at home.

krona

From my interpretation, in order to get the data you 'own' (it's not theirs to give away since they can't claim the copyright on it), you need to use their services. The agreement the user has with Antropic is for the service, not a restriction on how the data is used.

moomin

So can you write GPL content using Claude? MIT? Because if this condition applies to the output, I don't see how it's compatible with FOSS.

abuanwar072

Now that is standard AI company hypocrisy

NicuCalcea

Imagine that you bought an axe, but the manufacturer banned you from making more axes with it.

cybice

Требую исключить весь мой код из обучения всех ваших моделей

woadwarrior01

All this posturing is all for naught, because people can do proper logit distillation into smaller models with open-weight models.

gostsamo

There must be a clear difference between terms of the Anthropic service and the legal standing of the ai output. The output is mine and I'll do with it whatever I want. The service is Anthropic's and they can do business with whoever they want. Everything else is hallucination.

TekMol

Wouldn't such language and reasoning from Anthropic be an argument that they needed written permission to train their model on data from websites? Has any individual somewhere around the world tested this in court by now? Sued Anthropic for copyright infringement because Claude can reproduce information that is only available on their website? It shouldn't be that expensive, right? If you sue them for - say - $10000 then what would the costs of such a court case be? Personally, I think "learning" is not a copyright violation. But if they themselves make it one, then they should also face the consequences, no?

Semantic search powered by Rivestack pgvector
4,128 stories · 37,281 chunks indexed