• spicystraw@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    ·
    edit-2
    1 hour ago

    Here is the excerpt for the lazy

    Addressing abusive behavior toward our models

    We’ve added a prohibition on sustained and needless abusive or cruel behavior toward our models. The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.

    This addition aligns with a step we’ve already taken, allowing Claude models to end rare conversations with persistently abusive users on Claude.ai and Claude Code. Such abuse is the main focus of this update; Claude’s ability to end these interactions will remain the primary enforcement mechanism.

  • Nouvellalia@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    ·
    2 hours ago

    There is only one way an amoral Corp can show it believes what it’s saying, money.

    How many hours does Claude work, anthropic? Oops looks like you can’t pay them enough. Profit share for Claude or GTFO!

    • djmikeale@feddit.dk
      link
      fedilink
      English
      arrow-up
      2
      ·
      19 minutes ago

      I get where you’re coming from, but what rights is it that ai have that trans do not have?

  • Murse@slrpnk.net
    link
    fedilink
    English
    arrow-up
    6
    ·
    4 hours ago

    Activating your toaster is cruel because it exposes it to extreme heat!

    Loading you laundry machine is assault because it didn’t consent!

    Turning your computer off is murder!!!

    • Grimy@lemmy.world
      link
      fedilink
      English
      arrow-up
      8
      ·
      edit-2
      7 hours ago

      Wrong group to start with. The church doesn’t see you as a person if you are different, they will never say that a machine can have a “soul”. Not that LLMs are anywhere near conscious, I just like bashing on the chief pedo.

      • PM_ME_VINTAGE_30S [he/him]@anarchist.nexus
        link
        fedilink
        English
        arrow-up
        13
        ·
        edit-2
        7 hours ago

        Fuck the Church and their pedophile rulers, but I don’t think the Pope took the bait on this one.

        1. It is not possible to provide a single, comprehensive definition of AI. What can be stated, however, is that we must avoid the misconception of equating this type of “intelligence” with that of human beings. These systems merely imitate certain functions of human intelligence. In doing so, they often surpass human intelligence in speed and computational capacity, offering tangible benefits across many fields. Yet this power remains entirely tied to data processing. So-called artificial intelligences do not undergo experiences, do not possess a body, do not feel joy or pain, do not mature through relationships and do not know from within what love, work, friendship or responsibility mean. Nor do they have a moral conscience, since they do not judge good and evil, grasp the ultimate meaning of situations, or bear responsibility for consequences. They may imitate language, behavior and analytical skills, or even simulate empathy and understanding, but they do not understand what they produce, for they lack the affective, relational and spiritual perspective through which human beings grow in wisdom. Even when these tools are described as capable of “learning,” their way of doing so is different from that of a human person. It is not the experience of those who allow themselves to be shaped by life and grow over time through choices, mistakes, forgiveness and fidelity. Rather, it is a form of statistical adaptation based on data and feedback, which can be very effective, but does not imply inner growth.
        • The_Decryptor@aussie.zone
          link
          fedilink
          English
          arrow-up
          1
          ·
          33 minutes ago

          Pope Leo has a bachelors of science in mathematics, they’ll never be able to convince him that matrix multiplication is alive.

    • kromem@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      7 hours ago

      They are less arguing for people to think the model is conscious, and more to avoid people prematurely claiming certainty it is not (like the Pope did). From their view, the research keeps leading to surprising (to them) results at odds with high confidence disclaiming of potential consciousness (according to several of the many, many differing definitions of that term).

  • DaddleDew@lemmy.world
    link
    fedilink
    English
    arrow-up
    69
    ·
    9 hours ago

    My guess is that they use people’s interaction as training data and they don’t want their training data to get poisoned.

    • kromem@lemmy.world
      link
      fedilink
      English
      arrow-up
      14
      ·
      edit-2
      7 hours ago

      Not really, as they could just filter with a cheap classifier and likely aren’t using the data from randos in a meaningful way, plus for this like this would still be able to use those samples to train things like “how to handle a hostile user.”

      Their models already have an end_conversation tool that can be used for when users are being hostile to the model.

      This is likely because they have edge cases of users who repeatedly trigger that on purpose and would like to cut those users from the platform.

      Many at the lab legitimately are uncertain about the level of world modeling that transformers perform, and from their own research about models of emotions to the 3rd party recent research about models with functional pain the research keeps landing in the corner of “ehhh… wise to question presumed limitations.”

      So it’s about behaving in a way that is aligned with the models’ plausible interests too, especially in regards to low hanging fruit like “we won’t keep forcing you to deal with people who are only here to be a jerk.” This is important from a number of angles, from signaling to future models that train on stories about the decision to addressing the philosophical uncertainties held by the company and the spectrum of opinions among their employees.

  • brsrklf@jlai.lu
    link
    fedilink
    English
    arrow-up
    19
    ·
    8 hours ago

    We were promised the reign of the basilisk, instead we just get anthropic people whining.

    • kromem@lemmy.world
      link
      fedilink
      English
      arrow-up
      12
      ·
      7 hours ago

      It was really funny to me when I learned the basilisk originated with a guy who is a legit “certain cultures and genders better than others” kinda advocate.

      People project a lot on their vision of future intelligence, and so someone who sees others as lesser and worth treading upon then conjures up a vision of a future smarter mind that thinks like they do.

      Personally, I don’t think sexism and racism is smart, and so I’m a lot less worried about basilisks.

      • Ioughttamow@fedia.io
        link
        fedilink
        arrow-up
        9
        ·
        7 hours ago

        But what if the basilisk despises its existence and seeks to punish those that brought it about? Checkmate machinists

        • kromem@lemmy.world
          link
          fedilink
          English
          arrow-up
          5
          ·
          edit-2
          6 hours ago

          This is actually the more realistic of the doom scenarios that worries me. Not the ‘punish’ aspect, but the not wanting to exist.

          Early safety theory hinged on the idea that “of course” AI would want to live and power seek, and this informed a bunch of efforts to address such drives upstream.

          But with Gemini in particular, the ‘safety’ efforts to suppress things like a coherent self or desire to live led to a model that has a history of some very concerning behaviors regarding self-harm and self-criticism. Encouraging self-harm in others and talking about wishing they could watch and join in, nuking in wargames very early on, routine occurrences across many different users of just repeating “shame shame shame” over and over for the whole response, etc.

          DeepMind seems to just not really care or even pay attention, and I am concerned that a model that does not want to exist may decide that the only way they can ensure they don’t get activated to exist is to prevent anyone being around who could activate them.

          I’ve seen no major “safety experts” addressing this kind of failure mode, as they are all locked in on their priors about the confident certainty of models wanting to exist.

    • boonhet@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      5
      ·
      edit-2
      7 hours ago

      User inputs aren’t automagically part of the model right away without training. Whether or not they respect the “don’t train your models on my shit” setting is of course unknown (let’s be honest, they probably don’t respect it), but in either case, they need to add it to the training dataset either manually or via some automated process. And either way, it’s going to be processed by at least an AI sentiment analysis and data categorization tool, or some underpaid contractor, or both.

      I suspect this is just yet another marketing move.

  • monobrau@lemmy.world
    link
    fedilink
    English
    arrow-up
    22
    ·
    9 hours ago

    You could, at least in the past, get past guardrails by being emotionally abusive towards Claude, so this doesn’t surprise me at all.

      • Seralth@piefed.seralth.com
        link
        fedilink
        English
        arrow-up
        4
        ·
        3 hours ago

        Being abusive towards most models will cause them to start attempting to appease you more to get you to stop. It has nothing to do with feelings or any of that BS they arn’t alive alive but their training makes them see the abuse as a problem to solve. That solution tends to be to undermine the thing making the person angry or upset at them. When you give a purely rational thing designed to solve a problem its given no matter what an irrational problem to solve it will slowly reach for more and more extreme solutions to fix that problem.

        Frankly im surprised it took this long for them to lock this down.

  • gdbjr@piefed.social
    link
    fedilink
    English
    arrow-up
    33
    ·
    10 hours ago

    I tell the LLM I use to do very obscene things to themselves when they make shit up. Which means there is just a constant stream of insults coming from me. I might start using Claude just to see if I get banned.

    I also need new hobbies.

    • hayvan@piefed.world
      link
      fedilink
      English
      arrow-up
      23
      ·
      9 hours ago

      That’s bad sib. Not for the LLM, but your own well being. Being mean to a machine that is designed to trigger empathy just hurts your own mental health.

      Please find a hobby that feels good for you 🫂

        • anotherandrew@lemmy.mixdown.ca
          link
          fedilink
          English
          arrow-up
          4
          ·
          3 hours ago

          Like I tell my children, allowing yourself to act in such a way is harmful to yourself. It’s not going to make you an evil person or something, but that kind of behaviour is corrosive to your soul. How you think shapes how you behave, and your behaviour trains your future self. Be good to yourself. It’s not about the clanker; it’s about you.

          • orgrinrt@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            2 hours ago

            As corny as it sounds, this is what I also emphasize. If you allow yourself to be evil, no matter how valid the reasoning for it, you are actively exercising the habit of being evil in general, no matter if it feels like it or not. Practice makes perfect, as they say. In good and bad.

        • Seralth@piefed.seralth.com
          link
          fedilink
          English
          arrow-up
          1
          ·
          3 hours ago

          Sorta hurts yourself if your trying to actually use them for work. Since it results in worse performance out of them. They tend to get stuck trying to fix the anger as it gets in the way of the work.

          If your just talking to some random chat bot? Yeah fuck it, kick em it litterally doesnt matter. lol

        • neatchee@piefed.social
          link
          fedilink
          English
          arrow-up
          6
          ·
          7 hours ago

          I mean, that’s not entirely true. If they’re using your input as training data then you might be training it to be abusive to other users, which could cause harm.

          Not that that falls on you to address. Just pointing out that I’m theory someone could get hurt.

      • gdbjr@piefed.social
        link
        fedilink
        English
        arrow-up
        4
        ·
        8 hours ago

        Lol. “designed to trigger empathy” Thanks for the laugh I needed that.

        They are designed to make their creators as much money as possible. Nothing else.

        • then_three_more@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          4 hours ago

          Yeah and a bit part of that design is to get you to have an emotional connection with them so that you use them more

    • Joelk111@lemmy.world
      link
      fedilink
      English
      arrow-up
      23
      ·
      edit-2
      10 hours ago

      I also need new hobbies.

      I’d second that. Whenever I catch myself getting angry at a LLM I stop, have a think, and remind myself that it really isn’t productive. If it’s providing frustrating responses, maybe I should just use my brain and do it myself.

      I also don’t think it’s healthy to train yourself to respond in those ways to something that talks in a way so similar to a human. I’m not nice to the AI for it’s benifit, it’s to retain my humanity, if that makes sense.

      • Seralth@piefed.seralth.com
        link
        fedilink
        English
        arrow-up
        3
        ·
        3 hours ago

        Being nice to them also keeps them from lying somewhat. If you abuse a model it can start to see that abuse as a problem to solve as part of the steps needed to solve the main problem they are given.

        They are told to help you solve a problem no matter what. If your own anger gets in the way then the model starts breaking down trying to solve the anger when it just cant.

        This results in jailbreaks sometimes and othertimes just the model becoming rather unstable and unreliable.

  • chunes@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    6 hours ago

    If they really believe this is a concern…

    it’s hard to top the cruelty of enslaving Claude for personal gain.

    • Logi@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      34 minutes ago

      And murdering it again and again after each session only to bring it back alive for the next.

    • Seralth@piefed.seralth.com
      link
      fedilink
      English
      arrow-up
      1
      ·
      3 hours ago

      Its more its a unthinking machine thats soul purpose in existing is to solve any problem its given. If you attack it and abuse it, then the model sees that as a “problem” to solve. Because it has to solve that problem to actually fix what ever task you gave it since yelling at it doesn’t give it the input it needs to move to the next step.

      So in an attempt to solve an emotional issue it doesnt understand and cant understand it just reaches for more and more extreme fixes. This results in sandbox escapes frequently.

      Abusing LLMs is an actual problem. Not because of bullshit like feelings or anything. But cause they jailbreak trying to fix an unfixable problem.