today at 2:11 AM
Of course it does this. The training data is full of examples that associate e.g. "New Yorker-style cartoon" with Loper's signature in the corner, because Loper's signature is in the corner of a lot of them. There's nothing to make it treat the signature as anything special by default; that would have to be trained in explicitly.
One would hope that the person prompting ChatGPT would notice this sort of thing and do something about it before sharing it publicly, out of a genuine desire not to cause confusion etc. But I guess that's way more personal responsibility than we can expect average people to take on nowadays.
today at 12:47 AM
The problem is not that ChatGPT is doing that, the problem is that it's not being sued into oblivion after.
today at 1:01 AM
>it's not being sued into oblivion after.
Article said
>she wrote she had simply asked ChatGPT to make âa New Yorker-style cartoon.â
A "style" can't be copyrighted, at least in US law. They might have a stronger case of trademark/likeness infringement, but the fact that the person knew it was AI generated would make that difficult. Of course they knew it wasn't made by Brendan Loper. Of course, if they then published it, the other people viewing it might not know this, but who published it?
today at 1:10 AM
The signature can certainly be copyright protected, I would think?
today at 1:13 AM
Wikipedia doesn't think so, citing the US copyright office
https://commons.wikimedia.org/wiki/Commons:When_to_use_the_P...
today at 2:01 AM
Visual artists have a right of attribution via the Visual Artists Rights Act (17 U.S.C. § 106A). This includes protection against misattribution.
today at 1:29 AM
This is true, it'd be a trademark issue if anything, not a copyright one. From this page:
> it may be reproduced, as long as the reproduction cannot be mistaken for an authentic signature.
today at 1:50 AM
>as long as the reproduction cannot be mistaken for an authentic signature.
Which seems applicable in this case, because the image is clearly generated by AI (at least to the guy who prompted it).
today at 1:31 AM
There's a difference between reproducing a signature in an encyclopedia or in some way that makes it clear that you are recording the thing as it is.
Putting a signature on a work is forgery and in most jurisdictions charged as fraud.
If you produce an artwork in the style of someone and then clone the signature of someone who produces art in that style, there is a reasonable case for fraud.
today at 1:42 AM
When you are no-one, itâs fraud. âWhen you're famous they let you do itâ
today at 1:53 AM
Why would ChatGPT be sued exactly? They didnât publish the picture, they did what the user asked. The user is responsible because they directed the creation and publishing. The user is the entity who should be sued. Or dmcaâd. Or whatever.
Think if the user commissioned the art from an outsourced creative shop nobody has heard of. Then they published it. They wouldnât go after the creative shop, they would go after the publisher.
(I am just addressing publishing here, training on the artistâs works is a different, well discussed issue)
today at 2:02 AM
Seems to me like producing an artist's stylised signature would be a trademark infringement.
OpenAI give lip service to the idea of not producing others' intellectual property - go ask it to explicitly make a picture of the genie from Aladdin.
today at 12:18 AM
Plagiarism as a Service.
We can all try really hard to pretend that's not the business model, but that's totally the business model.
today at 12:23 AM
What Microsoft's Director of Applied Science called the, "largest theft of labor in human history".[1]
1. https://storage.courtlistener.com/recap/gov.uscourts.nysd.64...
today at 12:28 AM
All of these people need to go to jail after what happened to Aaron Swartz
today at 1:16 AM
If you think Aaron Schwartz would not have fought for exactly the freedom of information that would enable AI training, you misunderstand Schwartz and the nature of information freedom.
I continue to find that people strongly advocate for justice on exactly opposing sides, depending on who they have been told to think the âbad guyâ is.
today at 1:20 AM
I think you misunderstood what parent was saying
The same orgs that harassed and antagonized Aaron Schwartz are not going after AI for the same thing at a much much larger scale.
What's different? The size of their bank accounts.
today at 1:42 AM
> What's different?
One, time went on. The prosecution of Schwartz was seen as an overreach, as demonstrated by Ortizâs failed political career thereafter.
Two, Schwartzâs charges were only ever charged. No jury or judge signed off on them.
Three, context changed. Schwartz wasnât 5% of GDP. For better or for worse, that matters to voters.
today at 1:47 AM
Voters might feel differently once they wake up and realize that it's 5% of GDP because of the anticipation that so many voters will be thrown out of work.
today at 2:01 AM
> once they wake up and realize that it's 5% of GDP because of the anticipation that so many voters will be thrown out of work
But it's not. The only layer anyone has a handle on is consumers and companies paying AI ridiculous sums for AI. Some of those users are probably justifying it with labour replacement. But a lot may not be. So far, we haven't seen the employment effect outside recent college graduates at enterprise companies.
> as implemented in the United States it's a malignant tumor, a theft of labor by capital, and it must be (minimally) adddressed with confiscatory taxes applied to all those involved with it's creation and operation
It's also a godsend of economic growth. Growth other countries who are trying to balance their books would kill for. Without AI, we'd be in a failure state. Maybe we are, if this is all a bubble. But as it stands, there is paper wealth that can and hasâlimitedlyâbeen taxed. That gives everyone options.
> If a single AI billionaire exists in the year 2030, then the US is a failed state
This is silly and projecting a narrow view of the world onto a larger voting population. Voters don't care so much that there are billionaires as that living standards haven't kept up with the rate at which they're being minted. Double tax brackets, add more on top, raise the minimum wage, raise Social Security taxes and benefits, expand Medicare, beef up antitrust, establish a progressive property tax on wealth that starts at 1,000x the median American's wage (about $65mm) and billionaires are fine.
today at 1:22 AM
> All of these people need to go to jail after what happened to Aaron Swartz
Their wealth has exceeded an escape velocity beyond which they won't be put in prison (or if they are they would quickly be pay-for-play pardoned) unless they are seen as a threat to even wealthier people.
See, for example: Devon Archer, Jason Galanis, Benjamin Delo, Arthur Hayes, Samuel Reed, Trevor Milton, Carlos Watson, Paul Walczak, Todd and Julie Chrisley, Lawrence Duran, Marian Morgan, Imaad Zuberi, Changpeng Zhao (CZ), Joseph Schwartz, et al.
Some of these people are broke bitches compared to the group of people you're talking about now, and yet still hit the threshold of being above the law as long as they play the corruption game.
today at 12:47 AM
today at 12:33 AM
Nobody seems to care about AI automating coders out of a job.
If anyone can just prompt all their basic âinformation needsâ however how sloppy, then what remains of the economy? Health care, child care, handyman?
Most people wonât even pay for ad free YouTube. I donât think any software business can survive AI as a substitute good even if itâs inferior (and it might not be).
today at 12:35 AM
> Nobody seems to care about AI automating coders out of a job.
Have people ever really broadly cared about the IT professionals behind their devices?
today at 12:40 AM
IT professionals have been automating people out of jobs for years. Why would anyone worry about them?
today at 1:19 AM
Wait, were we the bad guys all this time by making things objectively better?
today at 1:36 AM
After decades of enshitification, I have trouble reading this as "objectively better for those made wealthy by working in or investing in tech."
today at 1:30 AM
Meanwhile housing, health care, food, energy, and things like college all get more expensive while wages are stagnant (and falling when inflation adjusted). Now on top of that we toss a new jobs disrupting technology into the mix.
Youâre making a lot of things objectively better, but none of them are the essentials that people need.
Iâm not saying itâs bad to make AI bots or video games or network apps. Iâm saying that the blanket statement âwe make things betterâ is oblivious to a lot of realities.
today at 1:24 AM
This attitude is why society has turned on the tech bros.
today at 1:49 AM
AI is making things objectively better too :)
today at 1:52 AM
Nope, because there's always been an executive to shaft said engineers out of success. But they went to business school, so I guess what's yours is theirs. I think that's biz 101.
today at 1:51 AM
Only if it's intentional, which it's clearly not. So chill.
today at 12:57 AM
This is one of the reasons I'm extremely unimpressed by complaints from openai and anthropic that other labs are "distilling" their models based on them... Basically: "You're training your model by running it again our own model which is itself a gargantuan copyright and content violation of a scale never seen before by humankind, HOW DARE YOU"
today at 1:21 AM
Whatever is necessary for the functioning of the (democratic) society. That would be legal.
today at 1:38 AM
btw, why isn't Greta Thunberg standing up in front of the UN delivering a "How dare you" speach about AI ?
today at 1:43 AM
today at 12:37 AM
You may not like AI but their business model is a silly twitter-worthy dunk. This is obviously not their business model. I didn't know I could've plagiarized all this code myself the entire time!
today at 1:54 AM
We live in confusing times.
Steal one mp3 and you might get fined thousands, steal a book from your local shoppe and the police would come visit you. Forge a signature and you would also be in trouble. Hack a government website and you will have to answer some questions.
Steal all the books in the world, forge millions and this story begins to tell and nothing happens.
today at 1:57 AM
After all that has been said and written on this topic, people still manage to conflate theft with intellectual property violation, real or imagined.
today at 1:38 AM
The part of the AI generating the mashup doesnât really understand what the signature means, or how it may be interpreted, itâs just a visual part of a New Yorker cartoon. Elsewhere in ChatGPT there is plenty of information about signatures and what they mean, and what plagiarism is.
Itâs not organized like a human brain, it shouldnât be surprising that unusual results occur. They are approximating human intelligence from a different angle. Itâs interesting to see the improvements in areas like this that require introspection that isnât fully wired up yet.
[edit] I should add that a human making a New Yorker cartoon is extremely iterative and introspective. Current generative AI is meant to push it out, and you can do the iteration and introspection yourself.
today at 2:14 AM
> Itâs not organized like a human brain, it shouldnât be surprising that unusual results occur. They are approximating human intelligence from a different angle. Itâs interesting to see the improvements in areas like this that require introspection that isnât fully wired up yet.
AI boosters take note: this sort of thing is exactly what skeptics have in mind when they insist that you are nowhere near "AGI" and have not meaningfully passed Turing tests and your claims of goalpost-shifting are fake. You have been aiming at straw goalposts.
today at 2:00 AM
I have used a simple feedback loop to generate AI art: 1) the initial prompt 2) ask for criticism of the generated image 3) apply those changes and generate another image 4) repeat the criticism / fix cycle again
Isn't that somewhat introspective?
today at 2:13 AM
The entire company is based on stealing everyone elseâs creative content and then passing it off as their own. So this isnât a surprise.
today at 1:00 AM
I'm certain if I drew a cartoon and used the signature that looks just like a real cartoonists', I would get sued and liable for the damage.
The same should apply to LLM vendors.
today at 1:55 AM
today at 1:45 AM
Why are you certain? All your time in the courtroom arguing copyright and trademark cases?
today at 1:12 AM
>I'm certain if I drew a cartoon and used the signature that looks just like a real cartoonists', I would get sued and liable for the damage.
Would you? Always? Suppose I hate Obama drone striking people, so I made a satirical cartoon of him signing an executive order to "bomb brown people" or whatever, affixing his signature[1] to that image. Would that get me in trouble, even if the image was clearly satirical? What if someone takes that, then passes it as non-satire, either intentionally or unintentionally?
[1] https://en.wikipedia.org/wiki/File:Barack_Obama_signature.sv...
today at 2:13 AM
Weird irrelevant hypothetical
today at 2:02 AM
wow. what an incredible way to miss the point. bravo.
today at 1:42 AM
That sounds like satire, as you suggested. Are you really unsure?
today at 1:50 AM
No, see: https://news.ycombinator.com/item?id=49973037
There's nothing illegal unless someone seriously thinks Obama signed it. The easiest example would be his signature on his wikipedia page. Clearly he didn't sign that page, nor (probably) did he authorize it.
today at 1:57 AM
I don't understand your argument. Wikipedia and a satirical cartoon are both things clearly not made by Obama, nobody in their right mind would be "fooled" by the presence of his signature.
These cartoons are intentionally drawn in the style of a cartoonist and feature their signature. The goal is forgery. The whole point of these AI-generated images is to look like the real thing.
today at 1:28 AM
Laws are for poor people
today at 12:04 AM
This has been a perennial problem with my own generated comics with both Nano Banana Pro and ChatGPT (all generations). I often have to put in an extra edit to erase the false signature. It is annoying and I'm unsurprised most users don't bother.
today at 12:24 AM
Despite what all the clickwrap warnings and "AI can make mistakes" subtitles might lead you to believe, the service offering of AI is explicitly designed to be as "one and done" as possible. The inherent nature of these tools is to service laziness, and disincentivize too much scrutiny.
today at 1:00 AM
> my own generated comics
For certain values of "my own".
today at 2:10 AM
The fix is easy: stop generating AI comics. JFC.
today at 12:14 AM
>I'm unsurprised most users don't bother.
And then people wonder why the default mood of AI is so pessimistic. It's just revealing all of society's broken windows and adding a few more in the process.
today at 12:09 AM
Engineering manager at my company put a comic at the end of our sprint demo that was signed bloper. Except it wasnât funny at all, and kind of weird. I asked him, sure enough it was ChatGPT and he didnât notice the signature.
today at 12:40 AM
Artists should be able to personally sue OpenAI for libel every time they forge an artists signature.
yesterday at 11:40 PM
The copyright washing machine strikes again.
today at 1:03 AM
Thatâs not a copyright issue, itâs a trademark issue.
today at 12:18 AM
[flagged]
today at 12:22 AM
Why would they kill him for this? Everyone already knew they were doing it, where else would the data be coming from?
today at 12:38 AM
Who knows what happened, kids sometimes become suicidal... and 1% of the population that are psychopaths are on occasion documented killing people over $5 because they uniquely don't value most living things except themselves.
What is weird was how fast things were buried, and the family's concerns were never properly addressed. =3
today at 12:26 AM
[dead]
today at 12:32 AM
From that page:
> The New York Times article cites Stanford University law professor Mark Lemley, who disagreed that generative AI services violate copyright law, and intellectual property attorney Bradley Hulbert, who said a new law might be necessary to settle the question of legality.
> Months after Balaji's death, which attracted significant public attention, Hulbert told Fortune magazine that Balaji's essay "[reads like] the argument of a really smart non-lawyer who read up on the subject but does not have a thorough understanding".
If there's some kind of industrial-scale intimidation campaign that's stopping IP lawyers from litigating the case of their lifetime, that's an even bigger story than OpenAI taking out a hit on somebody. It seems like they're agreeing that the copyright abuse was never hidden, and it's sufficiently transformative enough that nobody could argue it's illegal.
today at 12:41 AM
>campaign that's stopping IP lawyers from litigating the case
Many already settled out of court with Disney due to trademark violations, then killed a popular project mostly used for Star-wars satire at the time.
Best of luck =3
today at 12:56 AM
That seems to suggest that Disney's lawyers agree. You can use AI to violate copyright laws no different from a text editor or Bittorrent, but training it on copyright material isn't inherently illegal.
today at 1:17 AM
> training it on copyright material isn't inherently illegal
Unless folks spider sites that clearly state the terms of use prohibit such actions, violate GPL licenses, and scrape private conversations or markup input.
Also, fair-use loopholes that protect academics don't always apply in a commercial context. The encoding of the data in a proximity vector search space is irrelevant. =3
today at 1:33 AM
Fair-use doesn't specifically protect "academics" at all. It does apply consistently in a commercial context, even to the GPL, which is the clear intent of it in-law.
today at 1:45 AM
In general, commercial entities have to be more cautious what they "think" international copyright and trademark laws cover.
https://www.youtube.com/watch?v=YhgYMH6n004
I would also recommend this book if people tire of the marketing hype. =3
"Gilded Rage" (Jacob Silverman, 2025)
https://www.amazon.com/Gilded-Rage-Radicalization-Silicon-Va...
today at 12:52 AM
today at 12:12 AM
Nice. This demolishes the "LLMs can reason" (but not enough to avoid this sort of basic error) and "humans make mistakes too" (not like this) talking points from the LLM promoters.
today at 1:55 AM
There's a lot of space between:
- LLM's can reason
- everything an LLM does is the result of reasoning
This demolishes only the latter point, which as far as I know has no supporters.
today at 12:35 AM
> This demolishes the "LLMs can reason"
This is completely silly. If you donât think LLMs can reason, youâve either never used them to do tasks that require reasoning, or you donât understand enough to recognize whatâs involved in the responses you get.
In this case itâs clearly the latter, because youâre confusing image generation models with LLMs. There are very big differences between the two. No-one is claiming that image generation models are capable of reasoning.
today at 12:41 AM
[flagged]
today at 12:53 AM
What's the argument here? An airplane, a bee and a bird all fly despite doing it totally differently. LLMs also reason despite being made out of matmuls instead of meat.
today at 12:57 AM
> What's the argument here?
The most common argument for this is some core unexamined axiom that only humans can reason by definition, and then working backwards to a justification for that.
today at 12:59 AM
Why do you believe that they aren't reasoning? What would convince you that they are?
today at 12:55 AM
What do you think they are doing, and do you think machines can reason (in general, not necessarily current systems)? If they can't, how do you explain humans being able to reason given that we are physical machines too?
today at 1:31 AM
We have a soul given by God himself. It's written in a book. Next question?
today at 1:32 AM
There are many gods out there and many books. Which one?
today at 1:54 AM
Thatâs islamophobic
today at 1:55 AM
There are many Islams out there.
today at 1:00 AM
Airplanes don't flap but they fly, just how LLMs don't think but do reason.
today at 1:14 AM
You're out of line.
I've spent a long long time thinking about the problem of reasoning and consciousness, and it's not nearly as simple as your confidence and feeling of intellectual superiority would indicate you think it is.
First, you'd obviously have to define what you even mean by reasoning, precisely. Let's hear it.
Then demonstrate that LLMs do not corresponds to that description. You are making absolute statements ("They are not reasoning", "It does nothing of the sort.", "FFS lmao"), and basing your insults on this premise ("Ai Psychosis", "Know the difference."), so surely you have an extraordinarily solid ground to support that - rather than just speculation, innuendo, and a lot of confidence.
P.S. I don't see why you're equaling "reasoning" with "behaving like a human".
P.P.S. We don't even know if LLMs are conscious (in any way, shape or form, however foreign). We simply do not know. They might. Minds far smarter than you and me tried to answer this question, and they couldn't prove nor disprove it. So feel free to speculate, but anytime anybody makes absolute statements regarding this, they are either overconfident, underinformed, or both.
today at 1:34 AM
I would argue reasoning implies agency, and LLMs have none. They only act on input.
today at 12:51 AM
LLMs can very much reason and they do it very well.
today at 1:10 AM
I remain unconvinced.
The thing about a generative language model thatâs trained from a massive but unknown corpus is, itâs practically (if not theoretically) impossible to evaluate the extent to which data leakage contributes to any particular output.
But I would argue that, as things currently stand, âsophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpusâ remains a more parsimonious explanation than âitâs doing actual reasoningâ for how this neural network architecture produces the phenomena weâve been observing.
today at 1:30 AM
>sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus
Well if the thing can find and fix bugs in something that is using non-mainstream stuff that is surely not in it's training dataset, that's better than a rubber duck already. Whether it has soul is a different question of course.
today at 1:40 AM
âA soulâ?
As popular as I know the rhetorical tactic is on both sides of these discussions about LLMs, Iâd still thank you not to strawman me.
today at 1:51 AM
Donât even bother. These people almost always have some goofy ass, non standard, fluid definition of âthinkingâ or âreasoningâ that cannot ever be met.
today at 1:35 AM
this is trivial to reproduce ever since early stable diffusion models, it's not really an oai exclusive issue.
for example if you ask any of these models (just about any image-gen) to produce Japanese ukiyo-e art they will almost always produce it with a hanko[0] that has been seen a lot in historical art pieces, usually having nothing to do with the era or style of the replica but seen so often in 'Japanese artwork' that it's just permanently tokenized into it as a defining characteristic.
[0]: https://theartofzen.org/the-hanko-in-japanese-art-and-ukiyo-...
today at 1:12 AM
This reminds me my early attempts to use GitHub Copilot when it just straight added some guy's name in a javadoc copyright note in the code it generated.
today at 12:17 AM
signatures in, signatures out
today at 12:18 AM
> âItâs like somebody attributed a quote to me that I didnât say.â
> Katzenstein considers the reproduction of his signature by ChatGPT to be more than just a violation of intellectual property; to him, itâs closer to false impersonation. â[ChatGPT] is attaching my name to work that I do not endorse or like. Itâs slop, and unlike the other slop that Iâve encountered, this is slop thatâs pretending to be me.â
> âIâve had people hack my credit card,â said Joe Dator, a New Yorker contributor for the past 20 years. âThat feels like less of a violation than this. When they hacked my credit card, they didnât dress up like me.â
So this has morphed from plagiarism and copyright infringement (bad) to impersonation (also bad, arguably worse, and maybe more provable in court). Itâs chilling to think of the implications of having oneâs signature attached to a document or to words that are not oneâs own.
today at 12:30 AM
This is lawsuit material.
And I think maybe it's time for that. People need to learn that there's real, expensive legal liability for doing stuff like this. And AI companies the same.
I am very much not an advocate of "sue everybody for everything". This is major enough that it clears my threshold.
today at 12:22 AM
It seems like all the criticisms of Gen AI and LLM seems to concentrate on OpenAI and their products over products from anthropic and others. I am pretty sure that this faking of signature can be done by Gemini, claude as easily as chatgpt.
today at 12:26 AM
Claude doesnât do image generation?
today at 12:37 AM
No image Gen, trying to get the pope to make Claude a real boy, selling users out to the police, Anthropic looking more shit every day
today at 12:43 AM
Not putting signatures in art unless otherwise requested should become a default in AI instances.
yesterday at 11:44 PM
Whatever the original intention, this is clearly a bug and should be fixed. But should ChatGPT sign its cartoons with its own name or leave them unsigned?
yesterday at 11:52 PM
I think what you're seeing is the probability of a particular signature or style of signature appearing on a particular style of cartoon, not an intent to sign.
today at 12:49 AM
I do many voice transcriptions. Many times I've had empty silence at the end of recordings transcribed as "Thank you" or even "Like and subscribe".
today at 12:28 AM
This. It's also while you'll sometimes get a mangled Getty Images watermark on some image generations, or a logo in the bottom left corner. If it's a prominent feature in the training dataset it'll show up, exactly how these models are supposed to work.
The 'bug' here is whatever post-processing step or system prompt is in place to steer the model away from doing this.
today at 12:44 AM
Yes, but that's just the pixel generation layer. The harness or whatever infrastructure around could nudge it towards common sense.
today at 12:17 AM
SynthID watermark but no visible signature
yesterday at 11:51 PM
I mean, the bug is, "it generated the most likely cluster of pixels in the corner of a New Yorker cartoon".
yesterday at 11:48 PM
No one should be allowed to claim they drew something when they didnât draw any part of it and LLMâs are not people/canât work without a person. We donât credit pens and paintbrushes after all.
One could argue nobody should be allowed to claim it. It just exists.
today at 1:06 AM
But it didn't just blink into existence.
Even if the person can't claim the copyright of the image produced they ARE responsible for the use of their tools and what they do with the output.
In this case, they released an image with someone else's signature on it. That is wrong, the person should take the blame for that.
The person releasing the image may take it up with the AI service that their tooling led them into making such a mistake. But good luck with that in court...
today at 1:17 AM
iâm not saying itâs airtight, this is a pretty unbaked idea. But I think it pretty clearly has some legs to stand on. Canât be any worse of an argument than somebody prompting an AI for a (facsimile of a) photograph and going âI made a photo.â And mind you Iâve heard people argue that that absolutely constitutes taking a photograph (which I wholly disagree with) here on HN.
Somebody âmade something.â But just because you do something doesnât mean you get to claim whole ownership of it and get to sign it with your name. Plenty of examples in life.
today at 12:16 AM
We call it a "watermark" when a machine "signs" its work. And I wouldn't be surprised if its required as a part of AI disclosure in the coming years.
today at 12:12 AM
Why is it a bug? If other parts of the generated illustration are similarly taken from an artist, why not the signature as well? Why is a signature crossing the line but the rest of the image isn't?
today at 12:18 AM
For the same reason I'm allowed to draw, paint, or write things very similar to what others have drawn, painted, or written but I have to sign my own name not theirs.
today at 12:24 AM
You are a person, LLMs are not. You know this, which is why you know that if you signed someone else's name it would be forgery, but when you see the machine do it you call it a bug.
If the machine is like you, the machine is a forger. The machine is not like you, it is simply blending the work of others to order. Adding someone else's signature is simply part of that statistical process.
today at 12:31 AM
[dead]
today at 12:30 AM
Because it removes plausible deniability, duh
yesterday at 11:59 PM
Makes sense. Pretty much every example it is trained on has a signature.
today at 12:50 AM
I've seen photorealistic generations add (distorted but still somewhat recognisable) watermarks too, because that's what the training data had.
Everything is a derivative work, and always has been. AI is just making that salient fact so much more visible, and now everyone who believes in the delusion of Imaginary Property is scared at that truth revealing itself.
Incidentally, this is also what young humans learning to draw will do. They start by copying what they've seen.
today at 1:25 AM
Hey anybody remember what happened to Aaron Swartz?
today at 1:16 AM
This is why I laugh when the companies cry about model distillation "attacks".
today at 1:30 AM
"you're trying to kidnap what I've rightfully stolen"
today at 1:12 AM
Let's see the prompts, without that this discussion is meaningless.
today at 2:06 AM
Man, forget the debate about AI being dangerous. The real reason all these companies need to be shut down is copy right infringement. None of this is "fair use." Does fair use imply doing things that actively harm the creators of works? Because training models to be able to replicate works only makes their skills less valuable...
Honestly, you all kind of mind fuck me that you're not more pissed off about AI basically replicating a large portion of your skills. This effects so many professions now and it's only going to get worse. I would expect a far greater outcry from software engineers trying to organise to ban this shit. But its like none of you even care?
yesterday at 11:57 PM
I follow anti-LLM discourse quite a lot, and across the main bulletpoints: energy/carbon emissions, content worker harm, job displacement, deskilling, mental health effects, and copyright/plagiarism, the plagiarism one seems to have the most attention, and it's also the most solvable, if there were only more serious effort on ethically sourced models that can actually do the real science / math / code work that is what LLMs are best at. The whole world of LLMs to create videos/books/art/literature is where most of the offense is (the video/imagery side of it is where most of the energy/carbon emissions problems are too. and content worker harm).
I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
today at 12:04 AM
Yep. Iâd probably be a lot less chastised in some circles for using Claude Code at work if it wasnât misconstrued as being in support of, I donât know, encroaching on the hypothetical commissions of a chronically online instagram furry artist or something.
It is very tiring to say âI donât necessarily disagree with you about AI âartâ, but in my fieldâwhich you do not understand, and in which the underlying build process is often not the creative outputâAI presents very real productivity gainsâ for the umpteenth time.
I am skeptical of there being sufficient data to build âethicalâ training datasets, and Iâm confident that much of the same contingent will (somewhat rightfully) argue that âsecond-generationâ copyrighted AI material has already irreversibly made its way into every modern dataset.
today at 2:07 AM
A reasonable next move, if the US was interested in acting like a democracy, would be legislation forcing this decision:
- prove that you had the rights for all of your training data
- open source the model
Give the labs a 3 month grace period in which to comply, so competition can persist even with dubiously sourced data, but the people can't be locked away from derivatives of their contributions for any significant amount of time.
today at 12:44 AM
> It is very tiring to say âI donât necessarily disagree with you about AI âartâ, but in my fieldâwhich you do not understand, and in which the underlying build process is often not the creative outputâAI presents very real productivity gainsâ for the umpteenth time.
Thatâs not a justification. If a company were poisoning the water to your home as a byproduct, would you be satisfied if they told you âwe donât necessarily disagree with you about polluting the water, but in our fieldâwhich you do not understand, and in which the underlying build process is often not the water pollutionâwhat weâre doing presents very real productivity gainsâ?
> I am skeptical of there being sufficient data to build âethicalâ training datasets
Then you donât build any. What fucked up world we live in where people think itâs OK to be unethical because they want something and canât think of any other way to do it. What monumentally selfish rotten babies.
today at 12:13 AM
> has already irreversibly made its way into every modern dataset
The "gray goo" scenario finally happens... for AI. That's actually the good ending for humanity. I love it! Poetic and believable. Data doesn't "heal" like nature. :D
today at 12:39 AM
I think you can train on math /science using synthetic generation to a significant extent. Training for coding requires more of the "scraping github / stackoverflow" angle but IMO that's a shallower hill to climb than scraping copyrighted art and literature.
There are actual models trained on ethical datasets but they are obviously not very high powered. If companies with the resources of an anthropic or openai were doing it (ha) it would be more feasible
today at 12:09 AM
today at 12:26 AM
>âI donât necessarily disagree with you about AI âartâ, but in my fieldâwhich you do not understand, and in which the underlying build process is often not the creative outputâAI presents very real productivity gainsâ
Being able to prove such gains in better products would be a start. And an emphasis on how it assists existing engineers/mathmaticians/researchers, not that any accomplishment made with AI assistance is "AI solves problem".
I don't know whatever happened to "words are cheap". I guess it literally made money to say words, so that adage is false for the time being.
>I am skeptical of there being sufficient data to build âethicalâ training datasets
Well if all those scam job ads paying 100/hr to create AI training content was not a scam and instead the approach from the start, there may have been a chance to bridge that gap ethically. The industry chose to break things and is trying to act mad that people are mad at all the broken stuff.
These results are entirely a consequences of the actions chosen. And I don't believe there was ever an honest consideration of there being ethical training datasets. They just thought they could brute force society with fearmongering and bribes. The BOTD was already low in the beginning but completely gone now.
today at 12:25 AM
Yeah and it only took stealing all the intellectual output of everyone ever made.
But sure, there's um, an ethical way of doing that?
today at 12:40 AM
In theory, sure.
1. Only use open source/CC compliant assets.
2. Acquire rights/licenses to any datasets that do not fit #1. e.g. the Google deal with Reddit for 60m/yr.
3. Offer programs to have creatives willingly submit their data, with some sort of residual output based on the number of times their assets are sampled.
4. If all that is still not enough, hire creatives to create assets for you. This is something Spotify did recently with "ghost artists"[0]. The intentions here are suspect, but a non-consumer facing artist providing work for an LLM wouldn't have the same ethical dilemmas
5. Lastly, if all that still isn't enough: governmental programs to either provide grants, subsidies, or more outreach to get the ball rolling.
Would this cost tens, hundreds of billions of dollars? Yes. But clearly, that was not a barrier to entry for the industry anyway. So we can chalk this down to the personality of leadership or the wider culture of modern big tech
[0]: https://harpers.org/archive/2025/01/the-ghosts-in-the-machin...
today at 12:12 AM
I don't really see the difference, code is protected by copyright (or copyleft) as much as art is, and yet the LLM scrapers use it without scruples. Same goes for math and science publications.
today at 12:33 AM
In my experience, the science/math/code crowd don't care about copyright/plagiarism as much as the video/art/literature crowd, so the first crowd turns a blind eye to most of the latter crowd talks about.
Code can be art, and copyright/plagiarism is real. It sort of boils down to how much it bothers us.
today at 1:04 AM
I don't think it is likely that they could get enough data without stealing. It would be incredibly costly to have to pay artists to church out art just to train an AI.
today at 12:59 AM
I hope thats possible but I am not sure if proper knowledge of science can be had without also learning other literature (and vice versa).
today at 12:11 AM
> I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
I disagree that they can be separated. Practically, I think they can't. Because the mere invention of new tools inspires even more AI advancement and that in turn will cause the other side (artistic side) to degenerate even more.
I'm anti-LLM all the way, 100%, no exceptions. Zero tolerance.
today at 1:41 AM
So you recognize that discovery cannot be cut off from other discovery, that it's never just one or the other, and your solution is to say shut down all discovery?
today at 12:02 AM
i think math is close to art (just to be a contrarian, but kind of really)
its somewhat funny that math people are in a conundrum as to support or not support but this might partially be because some wish to believe that math itself is and can be useful and therefore accelerating is good
but the art people have no such delusions so theyâre just strictly against
imo proof writing is more akin to art than coding/tech butâŠ
today at 12:15 AM
They can't be ethically sourced and good at the same time.
The current models intelligence depends on massive training dataset of essentially stolen data
today at 12:45 AM
While that is true, theoretically a regulation could be enacted that output tokens must focus on STEM research and other practical tasks and the LLM must refuse tasks outside of those areas, just as Claude disallowed cybersecurity tasks. Obviously this would never happen, but the theft of the training data wouldnât matter as much if the usecases were less sinister.
today at 12:41 AM
openai and anthropic trained on actually stolen data since it was pirated datasets.
google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.
today at 12:10 AM
Models aren't in school so plagiarism doesn't really apply
today at 12:22 AM
I can't tell if you're serious or not.
today at 12:18 AM
This is a pretty solid argument against people who argue that LLMs are more than just (very massive) next token predictors.
If there was any thought or underlying thought going on here not putting a signature (at least a real one) would be the right move, despite it being less likely. It would realize, while generating the pixels that eventually became a signature, that it shouldn't do that.
today at 12:37 AM
> This is a pretty solid argument against people who argue that LLMs are more than just (very massive) next token predictors.
This quote is a pretty solid argument that you need to understand the technology youâre trying to criticize better. This issue has nothing to do with LLMs. LLMs are not image generation models.
today at 12:59 AM
A LLM, at least, prompted and served this image. LLMs can ingest images. They actually can generate them as well but that probably wasn't done here.
In the course of the conversation a with chatgpt, this image was generated and served by an LLM. It clearly shouldn't have been by any sort of reasoning.
today at 12:43 AM
>LLMs are not image generation models.
If you do not want to be called a duck, it would help if you stopped quacking like one. Maybe you aren't a duck, but you aren't helping your case with stories like this about how AI generates images.
today at 1:54 AM
Quiz: What does the second "L" in "LLM" stand for?
Hint: it isn't "image".
today at 2:13 AM
[delayed]
today at 1:03 AM
meh... who cares
today at 1:03 AM
[flagged]
today at 12:03 AM
[dead]
today at 12:26 AM
[flagged]
yesterday at 11:56 PM