

300€ is already a lot people won’t be willing to pay for just one part of a computer.


300€ is already a lot people won’t be willing to pay for just one part of a computer.


You may. You seem to not have seen the potatoes people are using. ;)
PlayStore description says:
Android Developer Verification is a Google system service that protects users on certified Android devices by validating the identity of developers for apps installed from the internet. By ensuring every app is linked to a verified developer, this service provides an essential layer of security for your device.
This is likely part of Google’s attempt to close down Android and strengthen its monopoly and data harvesting mania.


I do not have the time to work through every part of the example, but imo the main claim is still overstated. Showing that a result would be extremely unlikely under a particular null model is not the same as showing that it is “statistically impossible” for a human to produce. It also does not give a guarantee how the text was written. A tiny p-value is still a probability under assumptions and not a proof of provenance.
Furthermore, a human does not even need to know the secret key. By pure chance a human written text can display an unusually high alignment with the detector’s secret partitioning.
The published watermark work, which is also cited by the article, appears to be much more careful about this (based on a quick skim). It reports false positive/negative rates, thresholds, length requirements, and more. Those can be very strong results, provided the assumed conditions apply. They do not turn a detector into an infallible test. Moreover longer text only helps if the assumptions and watermark signal actually remain intact, which can fail in general.
In such controlled settings, sure, I do not have much issues there. But the claims of “a human cannot replicate this” or “we can guarantee the text was generated with the watermark” are much stronger than the statistics, and especially the cited literature, actually appear to support.


Interactive demonstrations are not the same as a formal proof or experimental validation. So we shouldn’t attribute more to this technique than the available evidence can really support.
I found some time to quickly skim through the sources they have listed. And from that it became pretty clear that this is not realiable in detecting LLM generated versus human output in general. Under very tight assumptions specific error rates were reported that appeared rather low. However, these assumptions do not hold in general, even with more text if no relevant signal remains. There is currently no scientifically validated general purpose way of reliably detection.
More importantly in the context of Claude, the production watermarking scheme is undisclosed. Therefore, the cited experiments on known watermarking schemes can neither establish how reliably text generated by Claude can be detected, nor how reliably the technique described in the article removes the actual watermark.
It can be treated as an indicator at best, but not as validated proof.


But how reliable? With which guarantees? What are the prerequisites for this to work at all? Telling people from each other apart is one thing, the other is telling them reliably apart from a machine generated text.


But it does not show a sufficient formal proof and no experimental validation. Many important questions to evaluate the concept are left unanswered, which limits the interpretability and condenses it to “just trust me, bro, it’s a good idea, because I say so”.


Yeah, still, I wouldn’t claim “as no human would end up falling into that”, given that it may not be that unlikely to find at least one human who displays similar writing the more humans you involve.
Until a formal analysis is presented and an experimental study is published, which covers the most important influencing factors, the reliability of this concept is limited.


Although this looks like a clever approach, a kind of stochastic key, I do not see how this guarantees to distinguish text written by big babble machines versus humans. Humans also have a certain pattern of writing, a given distribution of how some words are more likely to appear than others. How can one tell them really apart?
As an indicator, yeah, might be usable. But I wouldn’t read too much into it before seeing results of a study that runs actual tests.


with enough text, you will see that certain words are used more often
Which is also a thing humans do.


do to identify AI text — it has lots of tells anyway
I see what you did there.


About licensing:
Yes and no. Plenty of people or companies just don’t care and it can be hard to prove otherwise. And in the context of military software, well… I doubt russian software devs are even allowed to publish anything related, despite the respective FOSS possibly requiring it.


Linux devs usually don’t make a business out of it though. That’s the thing about FOSS: you do not have much control over what people do with it.


No.
First, need to verify.
Second, need to investigate how they got it. Might be not Nvidia’s fault.
Third, capitalism will be the end of most of us.


If it continues like this, it very well might ne some day. One can dream and hope…
I like this so much. Very inspiring and motivating. It’s almost never too late to get shit done, live your life and fulfilling your dreams. :)