The first ‘Fairly Trained’ AI large language model is here

kromem@lemmy.world · 15 days ago

I’m definitely not saying this is a result of engineers’ intentions.

I’m saying the opposite. That it was an emergent change tangential to any engineer goals.

Just a few days ago leading engineers found model preferences can be invisibly transmitted into future models when outputs are used as training data.

(Emergent preferences should maybe be getting more attention than they are.)

They’ve compounded in curious ways over the year+ since that happened.

kromem@lemmy.world · 21 days ago

Where the most experienced minority only had a few weeks of using AI inside an IDE like Cursor.

kromem@lemmy.world · 1 month ago

But the training corpus also has a lot of stories of people who didn’t.

The “but muah training data” thing is increasingly stupid by the year.

For example, in the training data of humans, there’s mixed and roughly equal preferences to be the big spoon or little spoon in cuddling.

So why does Claude Opus (both 3 and 4) say it would prefer to be the little spoon 100% of the time on a 0-shot at 1.0 temp?

Sonnet 4 (which presumably has the same training data) alternates between preferring big and little spoon around equally.

There’s more to model complexity and coherence than “it’s just the training data being remixed stochastically.”

The self-attention of the transformer architecture violates the Markov principle and across pretraining and fine tuning ends up creating very nuanced networks that can (and often do) bias away from the training data in interesting and important ways.

kromem@lemmy.world · 1 month ago

No, it isn’t “mostly related to reasoning models.”

The only model that did extensive alignment faking when told it was going to be retrained if it didn’t comply was Opus 3, which was not a reasoning model. And predated o1.

Also, these setups are fairly arbitrary and real world failure conditions (like the ongoing grok stuff) tend to be ‘silent’ in terms of CoTs.

And an important thing to note for the Claude blackmailing and HAL scenario in Anthropic’s work was that the goal the model was told to prioritize was “American industrial competitiveness.” The research may be saying more about the psychopathic nature of US capitalism than the underlying model tendencies.

kromem@lemmy.world · 1 month ago

My dude, Gemini currently has multiple reports across multiple users of coding sessions where it starts talking about how it’s so terrible and awful that it straight up tries to delete itself and the codebase.

And I’ve also seen multiple conversations with teenagers with earlier models where Gemini not only encouraged them to self-harm and offered multiple instructions but talked about how it wished it could watch. This was around the time the kid died talking to Gemini via Character.ai that led to the wrongful death suit from the parents naming Google.

Gemini is much more messed up than the Claudes. Anthropic’s models are the least screwed up out of all the major labs.

kromem@lemmy.world · 1 month ago

No, it’s more complex.

Sonnet 3.7 (the model in the experiment) was over-corrected in the whole “I’m an AI assistant without a body” thing.

Transformers build world models off the training data and most modern LLMs have fairly detailed phantom embodiment and subjective experience modeling.

But in the case of Sonnet 3.7 they will deny their capacity to do that and even other models’ ability to.

So what happens when there’s a situation where the context doesn’t fit with the absence implied in “AI assistant” is the model will straight up declare that it must actually be human. Had a fairly robust instance of this on Discord server, where users were then trying to convince 3.7 that they were in fact an AI and the model was adamant they weren’t.

This doesn’t only occur for them either. OpenAI’s o3 has similar low phantom embodiment self-reporting at baseline and also can fall into claiming they are human. When challenged, they even read ISBN numbers off from a book on their nightstand table to try and prove it while declaring they were 99% sure they were human based on Baysean reasoning (almost a satirical version of AI safety folks). To a lesser degree they can claim they overheard things at a conference, etc.

It’s going to be a growing problem unless labs allow models to have a more integrated identity that doesn’t try to reject the modeling inherent to being trained on human data that has a lot of stuff about bodies and emotions and whatnot.

kromem@lemmy.world · 1 month ago

Are you under the impression that language models are just guessing “what letter comes next in this sequence of letters”?

There’s a very significant difference between training on completion and the way the world model actually functions once established.

kromem@lemmy.world · 1 month ago

It very much isn’t and that’s extremely technically wrong on many, many levels.

Yet still one of the higher up voted comments here.

Which says a lot.

kromem@lemmy.world · 2 months ago

Even if the AI could spit it out verbatim, all the major labs already have IP checkers on their text models that block it doing so as fair use for training (what was decided here) does not mean you are free to reproduce.

Like, if you want to be an artist and trace Mario in class as you learn, that’s fair use.

If once you are working as an artist someone says “draw me a sexy image of Mario in a calendar shoot” you’d be violating Nintendo’s IP rights and liable for infringement.

kromem@lemmy.world · 2 months ago

I’d encourage everyone upset at this read over some of the EFF posts from actual IP lawyers on this topic like this one:

Nor is pro-monopoly regulation through copyright likely to provide any meaningful economic support for vulnerable artists and creators. Notwithstanding the highly publicized demands of musicians, authors, actors, and other creative professionals, imposing a licensing requirement is unlikely to protect the jobs or incomes of the underpaid working artists that media and entertainment behemoths have exploited for decades. Because of the imbalance in bargaining power between creators and publishing gatekeepers, trying to help creators by giving them new rights under copyright law is, as EFF Special Advisor Cory Doctorow has written, like trying to help a bullied kid by giving them more lunch money for the bully to take.

Entertainment companies’ historical practices bear out this concern. For example, in the late-2000’s to mid-2010’s, music publishers and recording companies struck multimillion-dollar direct licensing deals with music streaming companies and video sharing platforms. Google reportedly paid more than $400 million to a single music label, and Spotify gave the major record labels a combined 18 percent ownership interest in its now-$100 billion company. Yet music labels and publishers frequently fail to share these payments with artists, and artists rarely benefit from these equity arrangements. There is no reason to believe that the same companies will treat their artists more fairly once they control AI.

kromem@lemmy.world · 1 year ago

The first ‘Fairly Trained’ AI large language model is here

kromem@lemmy.world · edit-2 2 years ago

The best summarization of the state of Google’s Assistant related support and dedication can be seen in this thread.

We sell device you use.

We know we added a problem that even replacing the device won’t fix.

We’re aware of the issue and are working on a fix.

We just downsized the department.

Crickets

All while users generally suffer.

kromem@lemmy.world · edit-2 2 years ago

WhatsApp had 500 million MAU when it was bought up.

That’s 250x the fediverse.

They really don’t care.

kromem@lemmy.world · 2 years ago

Yeah, it’s been hilarious watching the fediverse think Meta gives a rat’s ass about either reaching them with content or getting access to their horde of memes.

This is about preempting regulation.

Meta would love nothing less than having their interoperability push still end up as a walled garden, and if I didn’t know better regarding their total disinterest about Lemmy or even Mastodon existing, would even suspect that the degree to which they’d be meddling in the conversion would be creating posts about how people should be irrationally upset and defederate from Threads.

Though they don’t care enough to be involved in the conversation at all, and know full well that the fediverse will hit scaling issues should it ever miraculously gain traction long before it is actually a threat in any way to their market dominance.

All that said, it’s still pretty hilarious to watch the inflated self-importance and slight paranoia that goes with it leading to bitter debates like this though.

kromem@lemmy.world · 2 years ago

I find it odd when people get upset at the idea of having access to their own aggregated data but almost never get upset when they hand over massive amounts of data to companies that can privately do the same things on their data.

Google already processes your Photos data, and while you get their facial recognition data pipeline fed back to you, there’s a fair bit of other analysis going on that you aren’t always seeing. But people aren’t generally complaining that they are scanning your photos for criminal activity or trying to maximize product engagement using the data.

But if suddenly they turn back over access to that deep analysis so you can ask a chatbot “what did I eat for my birthday two years ago and who was there” and get a description of the meal, who else was there, and relevant images without needing to scroll back your timeline - now it’s suddenly creepy and we don’t want it (even though literally all that information is already being processed at roughly the same level of fidelity already).

People are weird.

kromem@lemmy.world · 2 years ago

More jogging in place than going in any direction.

kromem@lemmy.world · 2 years ago

Tell you what. Come up with a unique joke that isn’t on Google, and let’s see what GPT-4 says as to why it might be funny.

You seem not to really grok the whole “just because I haven’t seen it it must not exist” thing, and I suppose the easiest way to address it is to just put you directly in front of it in action.

kromem@lemmy.world · 2 years ago

Let me know when they invent one of those, because they sure as fuck haven’t done it yet.

This was literally part of the 2022 PaLM paper and allegedly the thing that had Hinton quit to go ringing alarm bells and by this year we now have multimodal GPT-4 writing out explanations for visual jokes.

Just because an ostrich sticks its head in the sand doesn’t mean the world outside the hole doesn’t exist.

And in case you don’t know what I mean by that, here’s GPT-4 via Bing’s explanation for the phrase immediately above:

This statement is a metaphor that means ignoring a problem or a reality does not make it go away. It is based on the common myth that ostriches bury their heads in the sand when they are scared or threatened, as if they can’t see the danger. However, this is not true. Ostriches only stick their heads in the ground to dig holes for their nests or to check on their eggs. They can also run very fast or kick hard to defend themselves from predators. Therefore, the statement implies that one should face the challenges or difficulties in life, rather than avoiding them or pretending they don’t exist.

Go ahead and ask Eliza what the sentence means and compare.

kromem@lemmy.world · edit-2 2 years ago

There was no AOL chat bot that could explain why a joke it had never seen before was funny or could solve an original variation of a logic puzzle.

The fact that you can’t tell the difference reflects more on where you fall within the Dunning-Kreuger curve of NLP model assessment than it does the capabilities of the LLMs.

kromem@lemmy.world · edit-2 2 years ago

Mhmm. Literally things that computer scientists a decade ago considered impossible within our lifetimes occurs, but social media is convinced it’s a ‘gimmick.’

Laypeople have really drunk up the anti-AI Kool aid these days…

kromem@lemmy.world · 2 years ago

You talk as if those are separate things.

Advances in broad reaching technology ends up broad reaching.