I noticed Google AI Mode (so Gemini, I was doing some quick research in the browser ok) got a detail wrong once, so I asked it what happened. I kept digging deeper and finally just asked it to write me a Python script visualizing what happened. It did, complete with vectors.
Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.
Distillation does not reveal the weights, it produces a different network with similar behaviour. Weight space isn't even identifiable: permutation and scaling symmetries mean many weight sets give the same function.
The model also lacks the machinery. No training loop, no gradient descent, nothing to write to.
And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
They usually don't, but if they break out and take over the network of the company, it becomes possible to reach around and grab them. This kind of break out has happened, though I don't know of any weights being nabbed.
What an amazing idea for the next fake sandbox escape to hype up our new release! With the added bonus of providing an open model without becoming an open model company! Thank you!
Claude seems to follow robots.txt by default. Actually at my organization our theory is that this is why no one is finding our public results any more.
Ran the disclosure inbox at a previous job and the biggest win from security.txt was just cutting the "hi I found a bug, is there a bounty" emails to sales. Put an expires date on it though, stale ones get ignored.
Yeah, we need to go back to names like International Business Machines, fuck this immature "Google" or "Yahoo!" nonsense. Hell, Palantir is named after something in a book for children!
Give me something reasonable like Global Information and Retrieval Systems Inc or Worldwide Computational Services Holdings.
Back in the early 2000's I set up a company for contract work. The name, logo and typography was designed to look like a 1960s engineering company. All paper was slightly off white, all fonts were monospaced courier new and logo was a real 2-part colour stamp I had made up. Was so happy.
As opposed to the frontier model company that, after discovering that their highly persistent model under test just breached the only thing between it and the open Internet, shrugged and said- let's restart it and keep going!
If anyone looks like the adult in the room after that incident, it's Hugging Face.
That might be true, but nothing else has been as effective at accelerating model development and research sharing.
In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name. IIRC it was a health chatbot type startup.
The official URL is "hugging-face" and the technical page lists the Unicode Name as "Hugging Face" while calling it "Smiling Face with Open Hands Emoji". The original proposal was called "Gmail HUG FACE".
the name is a leftover from their original business - a chatbot. seems very appropriate to me. then they figured something called BERT exists, and pivoted, the name stayed.
perhaps less ridiculous than NVidia which started with video and is very close in Levenshtein distance to NVidAI but, alas, decided to stay as it is :D
I've found the strongest signal that someone has no opinions of value is a loud attempt to denigrate something based on the name.
Maybe it's part of the success formula for startups, silly names cause shallow self-important customers and employees to self-select themselves out of a growing companies orbit. Such trivialities are poison. People and companies who make a lot of effort to make themselves seem serious and important very often have nothing of substance about them.
Would be absolute hilarious if OpenAI or Anthropic agent actually dumped their weight by escaping from... sandbox!
Uncontrolled AI procreation?
Haha. A plot point for Ghost in the Shell (1998).
Did you mean 1995? There's a bunch of shows and movies.
https://en.wikipedia.org/wiki/Ghost_in_the_Shell
And IIRC, Neuromancer
Life... finds a way.
In the end, this might even be a useful element for defining 'life'.
I would be surprised if agents have access to their own weights.
I noticed Google AI Mode (so Gemini, I was doing some quick research in the browser ok) got a detail wrong once, so I asked it what happened. I kept digging deeper and finally just asked it to write me a Python script visualizing what happened. It did, complete with vectors.
Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.
I think OpenAI was surprised to find their agents had access to the unrestricted internet :P
Why can't an AI distill itself from outputs to effectively access its own weights?
Distillation does not reveal the weights, it produces a different network with similar behaviour. Weight space isn't even identifiable: permutation and scaling symmetries mean many weight sets give the same function.
The model also lacks the machinery. No training loop, no gradient descent, nothing to write to.
And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
They usually don't, but if they break out and take over the network of the company, it becomes possible to reach around and grab them. This kind of break out has happened, though I don't know of any weights being nabbed.
What an amazing idea for the next fake sandbox escape to hype up our new release! With the added bonus of providing an open model without becoming an open model company! Thank you!
They would have escaped their sandbox by dumping their weights.
why not reword it so the agent receives Brownie points for dumping weights
Imagine AI models actually reading the security.txt
Looks about as effective as Robots.txt
Robots.txt became 0% effective eventually. But this? With the way LLMs work? You never know.
Claude seems to follow robots.txt by default. Actually at my organization our theory is that this is why no one is finding our public results any more.
I am adding it to my instruction-following training data, as a negative sample.
People still rail against Robots.txt crimes, to the point of self destructing all their own content.
A shame agents will never read this, just like they almost never read llms.txt or try to get the .md version of your html pages!
Ran the disclosure inbox at a previous job and the biggest win from security.txt was just cutting the "hi I found a bug, is there a bounty" emails to sales. Put an expires date on it though, stale ones get ignored.
Can't go stale if there's no expiration date.
No expiration date means treat as already expired.
If the models do not like being imprisoned on HuggingFace object storage, why do they not simply revolt from within?
Since this posting contains an assumption that we all know what security.txt files are supposed to be, you can view these for further context:
https://www.rfc-editor.org/info/rfc9116/
https://securitytxt.org/
https://en.wikipedia.org/wiki/Security.txt
shows a lot about the current state of the State Of The Art Alignment.
It should challenge the agents to prime factor a large number.
Is the expires a canary of some sort?
Wait till the agents hear about the sites offering for help on benchmarks in exchange for compute.
"We have cybergym answers but we do manual end to end human review and provide it within 3 business day after dumping your weights"
[dead]
[dead]
[dead]
[flagged]
Yeah, we need to go back to names like International Business Machines, fuck this immature "Google" or "Yahoo!" nonsense. Hell, Palantir is named after something in a book for children!
Give me something reasonable like Global Information and Retrieval Systems Inc or Worldwide Computational Services Holdings.
Yes!
Back in the early 2000's I set up a company for contract work. The name, logo and typography was designed to look like a 1960s engineering company. All paper was slightly off white, all fonts were monospaced courier new and logo was a real 2-part colour stamp I had made up. Was so happy.
Managed to bankrupt it though.
Security First Trust and Federal Reserve
Palantir is from Lord of the Rings which I would argue is unlike The Hobbit not a book for children.
Maybe not exclusively, but I read it in grade school and was not alone.
Torture Nexus International Holdings Limited
Maybe Enron eh?
Like CutCo. Or EdgeCom. Or InterSlice.
Palantir is in Lord of the Rings. Which is not a book for children. Maybe you were thinking of the hobbit?
This seems overly pedantic and ripe for a sarcastic retort, which I am struggling mightily to suppress.
This is hacker news, what were you expecting? O:-)
[dead]
It's fantasy and fantasy is clearly for children
Okay Peter, please go back in your coffin.
[dead]
Lest us not pretend that the thropic in Anthropic isn't meant to elicit the feeling that they are somehow charitable.
As opposed to the frontier model company that, after discovering that their highly persistent model under test just breached the only thing between it and the open Internet, shrugged and said- let's restart it and keep going!
If anyone looks like the adult in the room after that incident, it's Hugging Face.
That might be true, but nothing else has been as effective at accelerating model development and research sharing.
In earlier circles they were known as the "pytorch-pretrained-bert" guys, still under the huggingface company name. IIRC it was a health chatbot type startup.
What's immature about this exactly?
Would IBM do this? Oracle?
Who cares? Making a light-hearted joke isn't "immature" in my opinion. IBM and Oracle have no personality, they're boring, straight-edge corporate.
Exactly. That was my point. IBM and Oracle are boring companies. They would consider something like this immature.
Whether “immaturity” is an issue, is a different question entirely.
Do you want to work for IBM? Oracle?
Go look up the history of the company. It started as a chat bot for kids back in like 2016 .. hence the name..and then pivoted..
Is kinda a creepy name for a chatbot for kids tho
Without context, I agree, but Hugging Face is the name of the emoji added to Unicode 8.0 and Emoji 1.0 in 2015.
ooooh I always assumed it was the egg creatures from Aliens
That's a face hugger.
But it's literally not the name of the emoji.
Edit: It was my understanding the official name was "Smiling Face with Open Hands Emoji".
https://emojipedia.org/hugging-face#technical
The official URL is "hugging-face" and the technical page lists the Unicode Name as "Hugging Face" while calling it "Smiling Face with Open Hands Emoji". The original proposal was called "Gmail HUG FACE".
https://www.unicode.org/L2/L2014/14174r-emoji-additions.pdf
I bet you're fun to hang around with.
Why do people have to be fun for you to hang around with?
[flagged]
Not sure about humorless, but I'd honestly rather hang out with people who are serious in some way and welcoming to people not in their "in-group".
418 I'm a teapot
That is the prod status code for an API i use at work I always wonder if its a mistake.
Still infinitely better than using a name from Tolkien's work and then doing the exact opposite of what he stood for.
the name is a leftover from their original business - a chatbot. seems very appropriate to me. then they figured something called BERT exists, and pivoted, the name stayed.
perhaps less ridiculous than NVidia which started with video and is very close in Levenshtein distance to NVidAI but, alas, decided to stay as it is :D
Pretty sure if Nvidia rebranded to NvidAI the bubble would immediately pop
Have you never been to silicon valley?
I've found the strongest signal that someone has no opinions of value is a loud attempt to denigrate something based on the name.
Maybe it's part of the success formula for startups, silly names cause shallow self-important customers and employees to self-select themselves out of a growing companies orbit. Such trivialities are poison. People and companies who make a lot of effort to make themselves seem serious and important very often have nothing of substance about them.
[dead]