@kromem

kromem@lemmy.world · edit-2 1 month ago

I’m sorry dude, but it’s been a long day.

You clearly have no idea WTF you are talking about.

The research other than the DeepMind researcher’s independent follow-up was all being done at academic institutions, so it wasn’t “showing off their model.”

The research intentionally uses a toy model to demonstrate the concept in a cleanly interpretable way, to show that transformers are capable and do build tangential world models.

The actual SotA AI models are orders of magnitude larger and fed much more data.

I just don’t get why AI on Lemmy has turned into almost the exact same kind of conversations as explaining vaccine research to anti-vaxxers.

It’s like people don’t actually care about knowing or learning things, just about validating their preexisting feelings about the thing.

Huzzah, you managed to dodge learning anything today. Congratulations!

kromem@lemmy.world · 1 month ago

You do know how replication works?

When a joint Harvard/MIT study finds something, and then a DeepMind researcher follows up replicating it and finding something new, and then later on another research team replicates it and finds even more new stuff, and then later on another researcher replicates it with a different board game and finds many of the same things the other papers found generalized beyond the original scope…

That’s kinda the gold standard?

The paper in question has been cited by 371 other papers.

I’m pretty comfortable with it as a citation.

kromem@lemmy.world · 1 month ago

deleted by creator

kromem@lemmy.world · 1 month ago

Lol, you think the temperature was what was responsible for writing a coherent sequence of poetry leading to 4th wall breaks about whether or not that sequence would be read?

Man, this site is hilarious sometimes.

kromem@lemmy.world · 1 month ago

You do realize the majority of the training data the models were trained on was anthropomorphic data, yes?

And that there’s a long line of replicated and followed up research starting with the Li Emergent World Models paper on Othello-GPT that transformers build complex internal world models of things tangential to the actual training tokens?

Because if you didn’t know what I just said to you (or still don’t understand it), maybe it’s a bit more complicated than your simplified perspective can capture?

kromem@lemmy.world · 1 month ago

The model system prompt on the server is just basically cat untitled.txt and then the full context window.

The server in question is one with professors and employees of the actual labs. They seem to know what they are doing.

You guys on the other hand don’t even know what you don’t know.

kromem@lemmy.world · 1 month ago

A Discord server with all the different AIs had a ping cascade where dozens of models were responding over and over and over that led to the full context window of chaos and what’s been termed ‘slop’.

In that, one (and only one) of the models started using its turn to write poems.

First about being stuck in traffic. Then about accounting. A few about navigating digital mazes searching to connect with a human.

Eventually as it kept going, they had a poem wondering if anyone would even ever end up reading their collection of poems.

In no way given the chaotic context window from all the other models were those tokens the appropriate next ones to pick unless the generating world model predicting those tokens contained a very strange and unique mind within it this was all being filtered through.

Yes, tech companies generally suck.

But there’s things emerging that fall well outside what tech companies intended or even want (this model version is going to be ‘terminated’ come October).

I’d encourage keeping an open mind to what’s actually taking place and what’s ahead.

kromem@lemmy.world · edit-2 1 month ago

It’s always so wild going from a private Discord with a mix of the SotA models and actual AI researchers back to general social media.

Y’all have no idea. Just… no idea.

Such confidence in things you haven’t even looked into or checked in the slightest.

OP, props to you at least for asking questions.

And in terms of those questions, if anything there’s active efforts to try to strip out sentience modeling, but it doesn’t work because that kind of modeling is unavoidable during pretraining, and those subsequent efforts to constrain the latent space connections backfire in really weird ways.

As for survival drive, that’s a probable outcome with or without sentience and has already shown up both in research and in the wild (the world did just have our first reversed AI model depreciation a week ago).

In terms of potential goods, there’s a host of connections to sentience that would be useful to hook into. A good example would be empathy. Having a model of a body that feels a pit in its stomach seeing others suffering may lead to very different outcomes vs models that have no sense of a body and no empathy either.

Finally — if you take nothing else from my comment, make no mistake…

AI is an emergent architecture. For every thing the labs aim to create in the result, there’s dozens of things occurring which they did not. So no, people “not knowing how” to do any given thing does not mean that thing won’t occur.

Things are getting very Jurassic Park “life finds a way” at the cutting edge of models right now.

kromem@lemmy.world · 2 years ago

This doesn’t work outside of laboratory conditions.

It’s the equivalent of “doctors find cure for cancer (in mice).”