janalsncm
yesterday at 10:39 PM
> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the âChatGPT momentâ wouldnât come until late 2022
I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.
(The only exception I will make is encoder-decoder models which now are often done by decoder-only.)
But what made it go mainstream was RL. RLHF at first, then other improvements like DPO that were less of a pain in the ass to set up. Adding diffusion on top of that would be an even bigger pain in the ass.
Before ChatGPT there really wasnât much of a concept of pre-training and post-training. It was all pre-training. Post training was what made the bots conversational and not just âcontinuing the thing you wrote to themâ.
So in short, diffusion never took off because it was just a more complicated way to generate tokens, and the real problem was getting tokens in the right distribution.
echelon
yesterday at 11:03 PM
> GPT2 was considered too dangerous to release
This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat.
Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies.
They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.
latentsea
today at 2:25 AM
What are your thoughts on the HuggingFace incident?
0xdeadbeefbabe
today at 3:47 AM
Optimal next token generation disguised as something more.
CamperBob2
today at 2:50 AM
"Drama in search of a moat" is hard to beat.
slashdave
today at 1:52 AM
Well, no. Without controls, a language model can drive sensitive individuals to violence or suicide. The idea of releasing a frontier model without RL is frightening based on what we have learned.
anon373839
today at 2:26 AM
> The idea of releasing a frontier model without RL is frightening
In case you were not aware, strong base models (no post training at all) have been available for quite some time now. Including ones that eclipse âscaryâ frontier models from even a year ago.
SV_BubbleTime
today at 3:08 AM
SoâŠ
Words out of a magic box on a computer canât make you kill yourself. Perhaps the people that would use that as encouragement are already mentally ill enough that it really doesnât matter what the trigger is?
On a more callus but fully serious note, evolution starts out physical, that the species that donât eat and breed as well, die. What happens when you remove that? When life is so safe that you basically wont stave, get eaten, catch a disease, donât really need to compete all that hard for scarce resources, etc? Do you think evolution just stops? Or perhaps does a social or mental evolution become the predominant differentiator for successful reproduction over time?
hodgehog11
today at 12:02 AM
I agree that this should be something that researchers reflect on. GPT-2 is one of the primary models to research on nowadays, and many recent developments have come from studying it as a test bench.
Imagine if CRISPR was considered "too dangerous" to publish because of the potential ethical ramifications, and that only a special few should be aware. It is utter self-righteousness, and it is shameful behaviour. The world cannot adjust itself to what it cannot see, so you risk greater catastrophe by keeping it secret.
The open dissemination of knowledge at every increment is the only way for society to truly deal with what is to come.
janalsncm
yesterday at 11:22 PM
My impression at the time was they were perhaps overly cautious but this was a bunch of researchers who wanted to self-regulate. Anthropic didnât exist, deepmind was also much less product-focused and relatively cautious. Chinese models werenât really a factor either.
Government regulation was really not in the picture either in 2020 or 2021 tbh. The government was still trying to beat a pandemic. Some of the Biden admin eventually wanted to but it wasnât very serious.
alightsoul
today at 1:11 AM
why are chinese models a consideration at all?
idiotsecant
today at 2:23 AM
Parent post is saying there was no community of peer models, just some people figuring it out as they went and with lots of slack to go slow if they wanted
alightsoul
today at 3:09 AM
Oh it's the typical china race thing, we can't let the antithesis of the us be better than the us etc. American exceptionalism
idiotsecant
today at 3:27 AM
The 'china' part is irrelevant. You're addicted to being mad. Parent post is just saying there were no other players