\

Guess which of these LLM outputs is watermarked

54 points - last Thursday at 2:03 PM

Source
  • Noumenon72

    last Thursday at 6:08 PM

    Please report success/failure after each test. Asking me to read and compare 30 writing samples to get any feedback at all means I won't finish. Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.

      • StilesCrisis

        today at 7:47 PM

        I did five, then gave up and just pressed A until I reached the end. I got 3/5 right.

          • neoncontrails

            today at 8:04 PM

            Exactly the same here. 4/5.

        • dozerly

          today at 7:03 PM

          Yea, I did two and then harrumphed in annoyance that I was expected to do all 10.

        • andai

          today at 8:05 PM

          > Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.

          Wouldn't this make it a worse measurement?

            • neoncontrails

              today at 10:02 PM

              Well, the alternative is a lot of us spammed A to get to the end, so the data quality is already horrendous.

          • stranded22

            today at 7:27 PM

            Yes. Did one - saw that I wouldn’t get feedback until I have completed all 10 (if at all) and noped out.

            • qarl2

              today at 9:24 PM

              I believe this is to test the theory that people can detect watermarked text.

              Teaching you how to identify watermarked text while the experiment is running would ruin the data.

                • ipaddr

                  today at 10:29 PM

                  Detecting watermarked text is hard. Detecting AI garbage text is easy but then classfiying the garbage further is beyond us.

                  • Phemist

                    today at 10:23 PM

                    Maybe detecting watermarked text is a skill to be attained. Not allowing proper feedback and training will not allow people to notice the difference on time, thus ruining the data?

                    Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.

            • Jowsey

              today at 10:27 PM

              Interestingly, it seems almost every set of three seems to follow a pattern: one passage of the three will have a key word or phrase swapped in the first sentence. That is, for every set of 3 passages, two will start with ~identical sentences, and one will have a key word or token changed.

              I caught onto this early and used it every time, and ended up getting 2/10, which is worse than random chance. I smell trickery!

              • bastawhiz

                today at 7:36 PM

                I've read that watermarking should in theory be impossible to detect except by the entity that watermarked it. Which is sensible, and I mostly understand at a high level.

                But what I don't know and don't understand is what happens if you watermark watermarked text. Does it test positive for both watermarks? Only the second? Indeterminate?

                Or maybe I'm misunderstanding. Can you tell that it's watermarked, but only the entity who put the watermark in place can test if it's theirs? My confusion about watermarking multiple times still stands, though.

                Regardless of what happens when you watermark multiple times, no matter the outcome, it weakens the watermark. Which, depending on the threat model, kind of makes it moot. I can't imagine a serious situation where a watermark can be weakened in any way and still be useful. Even "this came from an LLM" isn't a valid signal if you can just watermark ANY text through purely mechanical means.

                It's also not clear to me how this will affect mainstream LLMs. If all output text is watermarked, there MUST be an escape hatch. Otherwise, JSON schemas will break (or provide holes where unwatermarked text can be exfiltrated through MCP), "return this text exactly with no changes" will be impossible, and writing diffs will break.

                I feel like I must be missing something.

                  • mhitza

                    today at 8:21 PM

                    If you are refering in terms of the code of practice part of the EU AI Act

                    > I've read that watermarking should in theory be impossible to detect except by the entity that watermarked it

                    This is a carveout exception, for watermarking. In the spirit of those terms it should be machine identifiable.

                    In my opinion they should have thought better about this, paricularly for text, because in its current forms it is easy to lead next to a new "tamper-proof" requirement, which in practice is DRM. And we do not need more DRM.

                    For images, music there is metadata already where such information can be stored. And if end users are found using unlabeled AI their accounts could be ban from these platforms. Not something the social platforms might want, but it's a saner approach than trying to reinvent the secret printer dots on all generated media.

                    • red_admiral

                      today at 7:39 PM

                      Watermarking only applies where the AI has a free choice (https://www.anthropic.com/news/claude-text-watermark).

                        • 0xedwen

                          today at 8:04 PM

                          i think practical question is whether the detector survives ordinary transformations of the text. if for say i paraphrase a watermarked answer with another model, do we expect the original signal to disappear and the second model's signal to replace it?

                            • skybrian

                              today at 8:16 PM

                              There are multiple ways to paraphrase and the paraphrasing model is going to use its own random number generator whenever it thinks there's more than one possible choice. (Not really binary; the RNG will have more or less effect.)

                              So it will definitely be watermarked by the paraphrasing model. But the question is whether the original signal survives at all. There might be a weak signal that's detectable with enough text?

                              • billyp-rva

                                today at 8:16 PM

                                If its paraphrased to any significant degree I'd expect the original to not survive. The second model's watermark would of course be there regardless.

                            • throw310822

                              today at 7:52 PM

                              Does it still apply with zero entropy?

                                • red_admiral

                                  today at 8:40 PM

                                  No. Anthropic's example is completing the sentence "Isaac Newton's most famous work was called Principia ..." has only one correct answer, so nothing to watermark.

                          • today at 7:46 PM

                            • andy99

                              today at 8:18 PM

                              The watermarking is an inextricable part of the token generation, it just using a known pseudo random sequence for the sampling. It’s not a transform that can be applied later.

                                • rcxdude

                                  today at 9:59 PM

                                  This obviously stops working as soon as you don't have the entire context. To reliably detect a subset of the LLM's output you need to do something more sophisticated but also more invasive.

                          • fwlr

                            last Thursday at 4:41 PM

                            Utterly imperceptible, even when studied under the microscope in a way that LLM text very rarely is in practice.

                            It will be interesting to see whose concerns are assuaged (perhaps they genuinely though mistakenly believed it would degrade quality), and whose concerns are heightened (perhaps their real objection is that their AI-generated text will become detectable).

                              • red_admiral

                                today at 7:31 PM

                                If anyone notices degraded quality, that would imply they could break the crypto behind the watermark.

                                For an analogy, distinguishing AES ciphertext from random bits without the key would be counted as breaking AES (the more precise statement of this is called AEAD).

                                  • rcxdude

                                    today at 9:56 PM

                                    I'm not sure the watermark has been demonstrated to have that kind of property. Even if you use a CSRNG as the 'key' once you're feeding it through the token selection process it's going to risk re-introducing certain correlations.

                            • Jabbles

                              today at 10:22 PM

                              What is the meaning of the numbers?

                              > Weighted mean detector score 0.5307

                              Does it mean that that passage would be rated a 53% chance of being watermarked? So you would need a passage 10x as long to be reasonably sure of providence?

                              • lacker

                                today at 7:13 PM

                                This is like giving you three outputs from md5sum and asking you to guess for which one the input ended in a "q". There's no way to tell unless you break the RNG.

                                  • pllbnk

                                    today at 7:34 PM

                                    Yeah, it's a pointless exercise. I hope the author is just trolling given that he is knowledgeable in the field.

                                    > Here are three 64-character hex strings. Two are random. One is HMAC-SHA256(secret_key, "anthropic"). You don't have the key. Which one is the HMAC?

                                      • josh-sematic

                                        today at 7:48 PM

                                        I think the point is probably to help convince people that the watermarking doesn’t perceptibly impact quality, which is a concern some people have (whether well founded or not).

                                • reactordev

                                  today at 7:35 PM

                                  Wow I actually got a 7/10. It was hard to tell at first but there are signs that tipped me off to which one probably had a higher score out of the multiple choice.

                                    • petters

                                      today at 9:15 PM

                                      You likely just got lucky.

                                  • NotPractical

                                    last Thursday at 4:45 PM

                                    Could do with some context on how watermarking works. Objectively speaking it should be impossible to tell.

                                      • marcyb5st

                                        today at 7:16 PM

                                        My understanding is that watermarking in prose is basically a bias when sampling tokens. For a system that knows the average probability for each possible token in the LLM vocabulary it is possbile to quantify said bias given enough text.

                                        For a human that doesn't reason in tokens and therefore doesn't know anything about their probability distribution, it should be impossible to tell. Relying on fancy words/constructs within sentences should not give you any signal as well, since you don't know if the the prompt included instructions for that.

                                    • abathur

                                      today at 7:49 PM

                                      My sense of the concern here is that watermarking may somehow deprive someone or something of value regardless of whether or not they can tell, so I briefly pondered trying to rank these from best to worst and see if any set of those votes meaningfully deviated from ~average.

                                      That said, I read the first triple and found all three tortured enough that I can't be bothered with the rest.

                                      Call me persuaded, I guess.

                                      • Lerc

                                        today at 7:29 PM

                                        I was never going to do very well on this. My ADHD was itching after the third one. I suspect it would have been sooner but I had a bit of extra focus from the suprise that it selected an answer for the first question when I tried to scroll.

                                        To avoid that on the following questions I just held my finger on my phone to avoid a click. That eventually selected some text, and I instinctively tapped to deselect. That triggered another random pick, then I just tapped through to the end because I was fed up.

                                        • rrr_oh_man

                                          today at 7:37 PM

                                          It feels all of them are terribly written. I don't know why.

                                            • StilesCrisis

                                              today at 7:50 PM

                                              Because it's AI slop? Not that surprising.

                                          • arcwhite

                                            last Thursday at 2:19 PM

                                            Interesting, I did very badly, 3/10!

                                              • bastawhiz

                                                today at 7:24 PM

                                                Same, doing worse than random chance seems like an interesting signal though, but I'm not sure what it's a signal of.

                                                  • DHowett

                                                    today at 7:33 PM

                                                    Only one third of the options at each stage are watermarked, so 3/10 seems well within random chance.

                                            • madarcho

                                              last Thursday at 4:53 PM

                                              If SynthID is a google technology, then this is likely just us training their ai again, captcha all over again.

                                              • smallerize

                                                today at 7:29 PM

                                                Google's SynthID page says they can watermark text, but it also says that it can only detect the watermark on "image, video or audio". Does that mean that the text watermarks can't actually be used as watermarks?

                                                • aizk

                                                  today at 10:12 PM

                                                  I had a moment I thought was concrete watermarking the other day. Claude wrote the sentence... "since the compute buffer estimate has some sl..." Now you'd think the right word would be slack, but Claude wrote... slop? Which does seem off but, those two letters could be tokens very close in probability.

                                                  • elikoga

                                                    today at 12:03 AM

                                                    I disliked the fact that the experiment only covered prose, which my eyes glossed over and made me actually do random entries to pass on and see the results. I'd love to see it on a more accurate output distribution like commented code

                                                      • gjm11

                                                        today at 7:00 PM

                                                        I don't think anyone is, or plans to be, watermarking AI-generated code as opposed to text.

                                                        [EDITED to add:] As pointed out by a helpful comment below, I was misremembering: Anthropic do apply their watermarking to code, they just say that it will have negligible impact on the actual code (because there's generally less scope for variation in that) but e.g. it will have its usual effect on comments in the code.

                                                          • raincole

                                                            today at 7:26 PM

                                                            https://www.anthropic.com/news/claude-text-watermark

                                                            > code—which in very many cases has to be exact—has generally less watermarking than some other forms of text.

                                                            Generally less watermarking. Not no watermarking.

                                                            • demibabs

                                                              today at 7:20 PM

                                                              What do you mean? The watermarking applies to all text outputs, including code.

                                                              It’s just much less effective since code is low entropy.

                                                              • red_admiral

                                                                today at 7:18 PM

                                                                Why not? It would help with a lot of potential legal issues.

                                                        • jdw64

                                                          today at 10:11 PM

                                                          8/10. It was harder to distinguish than I expected. If they had applied something like a humanizer skill, it probably would have been nearly impossible to tell.

                                                          • smikhanov

                                                            today at 7:48 PM

                                                            It takes a lot of patience to read this much slop voluntarily.

                                                            • today at 8:19 PM

                                                              • kshmir

                                                                today at 6:51 PM

                                                                Thought there were only 2 options!

                                                                  • foundry27

                                                                    today at 7:36 PM

                                                                    Me too lol.

                                                                    It was only at question #8 that I realized there was a third option, and while I’d love to say that accounts for how I got a 1/10, after reviewing the third options I doubt it would’ve made a difference.

                                                                • jibal

                                                                  today at 10:00 PM

                                                                  This is stupid --- way too long and wordy. Make your test worth taking. And even then it's theoretically impossible to detect the watermark so what even is the point? If it's to check whether the watermarking actually has that property, this is not at all a reliable way to do that.

                                                                  • AiToolsGem

                                                                    today at 10:27 PM

                                                                    [flagged]

                                                                    • MagicMoonlight

                                                                      today at 9:27 PM

                                                                      [dead]

                                                                      • lolokbro

                                                                        today at 7:17 PM

                                                                        [flagged]