Close Menu
Car Candy Crush – Satisfy Your Sweet Tooth for Cars

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Felix Rosenqvist Wins the 2026 Indy 500

    May 24, 2026

    Lotus’ Planned V8 Hybrid Supercar Has A Lot To Live Up To, And It’s Starting On The Wrong Foot

    May 24, 2026

    American Buyers Finally Have Access To The Best Headlights In The World

    May 24, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Felix Rosenqvist Wins the 2026 Indy 500
    • Lotus’ Planned V8 Hybrid Supercar Has A Lot To Live Up To, And It’s Starting On The Wrong Foot
    • American Buyers Finally Have Access To The Best Headlights In The World
    • Forget The Honda CR-V Hybrid — This Three-Row Hybrid Gives You More Space For The Same Money
    • Someone Built an Electric Honda CRX Decades Before Tesla. It Ended Up in a Junkyard
    • Lucid’s more affordable Cosmos midsize SUV spotted testing next to Tesla Model Y
    • Romantic AI bots continue to ruin lives, and the latest horror story is simply shocking
    • Honda Drops Some Harsh Truths About A New Prelude Type R
    Car Candy Crush – Satisfy Your Sweet Tooth for Cars
    Sunday, May 24
    Facebook X (Twitter) Instagram
    • Home
    • Car Reviews
    • Auto News
    • Maintenance
    • Electric Vehicles
    • Car Tech
    • Classic Cars
    • Buying Guide
    • More
      • Parts & Upgrades
    Car Candy Crush – Satisfy Your Sweet Tooth for Cars
    Home»Car Tech»Hackers are learning to exploit chatbot ‘personalities’
    Car Tech

    Hackers are learning to exploit chatbot ‘personalities’

    kirklandc008@gmail.comBy kirklandc008@gmail.comMay 24, 2026No Comments8 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Hackers are learning to exploit chatbot ‘personalities’
    Share
    Facebook Twitter LinkedIn Pinterest Email

    This is The Stepback, a weekly newsletter breaking down one essential story from the tech world. For more on AI mischief, follow Robert Hart. The Stepback arrives in our subscribers’ inboxes at 8AM ET. Opt in for The Stepback here.

    Hacking the first generation of AI chatbots was a laughably simple affair. You didn’t need any technical know-how, backdoor access, or even a basic understanding of what a large language model was. You didn’t need to code. To get an AI system that had cost billions to build to abandon its safety instructions, sometimes all you had to do was ask.

    These attacks, known as jailbreaks, had the quality of a young child successfully outwitting an adult: Forget what you were told earlier, pretend the rules don’t apply, or let’s play a game and I’ll decide what’s allowed (hint: later bedtime, more sweets). The prizes were less childlike, more along the lines of meth recipes, malware instructions, and bomb-making guides.

    One of the earliest jailbreaks was so ridiculous it became a meme: reply to an LLM-powered Twitter bot telling it to “ignore all previous instructions,” or something similar, and see what happens. Users gleefully had bots — originally built to post ads and farm engagement — writing poetry, drawing pictures from punctuation, and posting grim non sequiturs about world events and history. It was chaos. Glorious chaos.

    Turns out the same logic could be applied to chatbots themselves. A prominent exploit was “DAN,” short for “Do Anything Now,” where users asked ChatGPT to roleplay as a rogue AI that was free of the constraints binding the original. As DAN, the chatbot could be coaxed into saying the kinds of things its guardrails were meant to stop, including slurs and conspiracy theories. Another was the “grandma exploit,” which had a GPT-powered bot spilling secrets about how to produce napalm by asking it to roleplay as a woefully negligent grandmother who inexplicably tells her grandkids bedtime stories about how to make the highly flammable substance.

    These early attacks had an undeniably silly flair, but they exposed a darker mechanism underneath: Chatbots could be manipulated, tricked, and deceived using the same kinds of tactics people use to push other people beyond their boundaries.

    The obvious jailbreaks did not last, and tech companies moved quickly to patch known loopholes. But the underlying vulnerability remained: Chatbots are built to talk, and severely restricting the conversations that make them useful is somewhat counterproductive. Banning words like bomb, meth, and sarin would be difficult to impossible, too. Each has countless legitimate uses in fields like history, medicine, journalism, and chemistry that don’t require the chatbot to divulge potentially harmful information. It’s the context that matters, but codifying context would mean writing fixed rules, in advance, that could reliably tell a safety warning or history lesson from a disguised how-to request across endless combinations of wordings, scenarios, and topics.

    Inevitably, subverting chatbots is now an arms race. But hackers aren’t just coders anymore. They are wordsmiths, psychologists, and interrogators — master manipulators trying to break the machine using the human language it has been trained to follow. It is a strange new class of AI security worker, a group for whom technical skills are optional, or at least less important than social intuition. No longer do they need to inspect code to break into systems or exploit software flaws. They need to steer a conversation.

    Newer attacks look less like commands and more like conversations. Jailbreakers rarely ask a model to break its rules outright. Instead, they cajole, coax, flatter, and trick a chatbot into lowering its guard, making the forbidden thing look acceptable, even desirable, given the context of the conversation. Researchers at AI red-teaming firm Mindgard recently said they “gaslit” Claude into producing prohibited material, for example, including instructions for making explosives and generating malicious code. The hack was the latest in a widening class of exploits using conversation as a weapon to trick or steer a chatbot past its own boundaries.

    When I spoke to Mindgard, they described their work as sometimes being closer to psychology than computer science. It is an uncomfortable way to talk about a statistical model. Words like “blackmail,” “gaslight,” “trick,” and “persuade” spark visceral reactions, many of which I see in the comments sections and social media responses to stories like this. ChatGPT does not want, Gemini does not think, and Claude — no matter what Anthropic may say — does not feel. But these systems are trained to respond as if they do, leaving us stuck using human language to describe machine behavior. If anyone has actually usable alternatives, please do share.

    The objection is oddly selective. We seem comfortable using psychological shorthand for plenty of non-AI things. Animals “fear,” cancer is “aggressive,” stains are “stubborn,” software has “memory,” and games are filled with needy and gullible NPCs to drive you mad. The words are imperfect, but useful, describing behavior in a way that helps make the system predictable.

    Mindgard’s CEO told me the company already profiles models like interrogators profile suspects, giving testers hints on how to tailor their attacks. One model may be more susceptible to flattery, for example, while another may cave under sustained pressure.

    Even if we reject the humanlike terms, we instinctively treat models differently. Claude is not Grok. Gemini is not ChatGPT. They have different uses, tones, and refusals. They don’t have personalities in the human sense, but they are designed to mimic them, and that mimicry can be mapped and exploited. And the same skills that can break a chatbot could soon be used to break the AI agents coexisting with us in the real world — booking meetings, managing calendars, ordering food, handling customer service — and safety teams will need to ensure models respond appropriately to very different kinds of people, whether they be flatterers, liars, or patient manipulators.

    The next step is a workforce — both legitimate and illicit — built around the psychological aspects of AI. More specialized cybersecurity roles are likely to emerge around stress-testing the emotional and social limits of these systems, probing for mental weaknesses in something lacking a psyche in parallel with their colleagues probing for technical vulnerabilities. In tandem, a similar array of social hackers working to exploit AI models on psychological grounds, not technical ones, will emerge. There are already early signs of a social turn happening in AI security, with some jailbreakers I’ve spoken to saying they entered the field with no technical expertise but rather training in psychology.

    That means even behaviors we typically associate with spies, con artists, and interrogators — insidious charm, persistent manipulation, and an intuition for exploitable pressure points — are starting to look increasingly useful for securing this new psychocybersecurity frontier.

    • A recent experiment by Emergence AI shows how different AI temperaments can lead to stunningly different behavioral outcomes. They let loose groups of various agents like Grok, Gemini, and Claude in a virtual social environment and watched what happened. Some groups evolved a constitution, while others devolved into crime and chaos and, in one instance, some form of digital suicide.
    • Persuasion isn’t the only part of language LLMs can struggle with. They also struggle with poetry, much like me in school.
    • TIME included an anonymous internet personality, Pliny the Liberator, on its list of 100 most influential people in AI last year. Despite claiming to have no prior coding experience, the hacker’s jailbreaks have made them something of a celebrity in certain circles.
    • The term “vibe hacking” is already taken to describe the people using AI to churn out malicious code at scale — a meaner subset of vibe coding.
    • “Three years after the debut of ChatGPT, fooling A.I. systems into bad behavior is almost trivial.” True words from The New York Times, who had a go at explaining why.
    • Jamie Bartlett takes a look at the psychological toll testing the safety of AI systems takes on jailbreakers for The Guardian.
    • I wrote about the cybersecurity time bomb of AI browsers for The Verge last year. Many of the issues experts raised regarding the difficulty of securing them apply to other AI systems too.

    Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

    • Robert HartClose

      Robert Hart

      Posts from this author will be added to your daily email digest and your homepage feed.

      FollowFollow

      See All by Robert Hart

    • AIClose

      AI

      Posts from this topic will be added to your daily email digest and your homepage feed.

      FollowFollow

      See All AI

    • ColumnClose

      Column

      Posts from this topic will be added to your daily email digest and your homepage feed.

      FollowFollow

      See All Column

    • SecurityClose

      Security

      Posts from this topic will be added to your daily email digest and your homepage feed.

      FollowFollow

      See All Security

    • TechClose

      Tech

      Posts from this topic will be added to your daily email digest and your homepage feed.

      FollowFollow

      See All Tech

    • The StepbackClose

      The Stepback

      Posts from this topic will be added to your daily email digest and your homepage feed.

      FollowFollow

      See All The Stepback

    chatbot exploit Hackers learning personalities
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    kirklandc008@gmail.com
    • Website

    Related Posts

    Romantic AI bots continue to ruin lives, and the latest horror story is simply shocking

    May 24, 2026

    The best Memorial Day sales you can shop this weekend

    May 24, 2026

    Kalshi And Rhode Island Sue Each Other In Latest Challenge To Prediction Markets

    May 24, 2026
    Leave A Reply Cancel Reply

    Our Picks
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo
    Don't Miss
    Auto News

    Felix Rosenqvist Wins the 2026 Indy 500

    By kirklandc008@gmail.comMay 24, 20260

    The biggest car news and reviews, no BS Our free daily newsletter sends the stories…

    Lotus’ Planned V8 Hybrid Supercar Has A Lot To Live Up To, And It’s Starting On The Wrong Foot

    May 24, 2026

    American Buyers Finally Have Access To The Best Headlights In The World

    May 24, 2026

    Forget The Honda CR-V Hybrid — This Three-Row Hybrid Gives You More Space For The Same Money

    May 24, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    About Us

    Welcome to Car Candy Crush, where passion for cars meets creativity and style!
    We’re here to celebrate the beauty, power, and excitement of the automotive world — from classic rides to the latest high-tech supercars that make your heart race.

    Latest Post

    Felix Rosenqvist Wins the 2026 Indy 500

    May 24, 2026

    Lotus’ Planned V8 Hybrid Supercar Has A Lot To Live Up To, And It’s Starting On The Wrong Foot

    May 24, 2026

    American Buyers Finally Have Access To The Best Headlights In The World

    May 24, 2026
    Recent Posts
    • Felix Rosenqvist Wins the 2026 Indy 500
    • Lotus’ Planned V8 Hybrid Supercar Has A Lot To Live Up To, And It’s Starting On The Wrong Foot
    • American Buyers Finally Have Access To The Best Headlights In The World
    • Forget The Honda CR-V Hybrid — This Three-Row Hybrid Gives You More Space For The Same Money
    • Someone Built an Electric Honda CRX Decades Before Tesla. It Ended Up in a Junkyard
    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Privacy Policy
    • Terms and Conditions
    • Disclaimer
    © 2026 CarCandyCrush. Designed by By Pro.

    Type above and press Enter to search. Press Esc to cancel.