Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Nscale buys Anyscale as it seeks to own more of the AI compute stack

    July 31, 2026

    Daisy One Review: Comfy and Tactile Headphones

    July 31, 2026

    How the 60 Minutes Editorial Process Is Supposed to Work

    July 31, 2026
    Facebook Twitter Instagram
    • Tech
    • Gadgets
    • Spotlight
    • Gaming
    Facebook Twitter Instagram
    iGadgets TechiGadgets Tech
    Subscribe
    • Home
    • Gadgets
    • Insights
    • Apps

      Nscale buys Anyscale as it seeks to own more of the AI compute stack

      July 31, 2026

      Netflix lands global streaming deal for ‘The Walking Dead’

      July 31, 2026

      Meta says AI is making it easier to build new apps — and more are coming

      July 31, 2026

      When will fusion power startup Commonwealth Fusion Systems go public?

      July 31, 2026

      Spotify launches ‘User Notes’ to let users add memories to songs

      July 31, 2026
    • Gear
    • Mobiles
      1. Tech
      2. Gadgets
      3. Insights
      4. View All

      Daisy One Review: Comfy and Tactile Headphones

      July 31, 2026

      The New Defcon Badges Pack a Unique Open Source Chip That Doubles as a Security Key

      July 31, 2026

      The World Is Too Hot. El Niño Is Partly to Blame

      July 31, 2026

      SelectBlinds Promo Codes & Coupons: Save on Custom Window Treatments

      July 31, 2026

      March Update May Have Weakened The Haptics For Pixel 6 Users

      April 2, 2022

      Project 'Diamond' Is The Galaxy S23, Not A Rollable Smartphone

      April 2, 2022

      The At A Glance Widget Is More Useful After March Update

      April 2, 2022

      Pre-Order The OnePlus 10 Pro For Just $1 In The US

      April 2, 2022

      Motorola Edge+ Review: It Checks A Lot Of Boxes

      April 2, 2022

      This Smartphone Concept Design Is Different… In A Good Way

      April 2, 2022

      Twitter Just Made Searching Your Direct Messages Better

      April 2, 2022

      That Netflix Price Hike Is Starting To Take Place

      April 2, 2022

      Latest Huawei Mobiles P50 and P50 Pro Feature Kirin Chips

      January 15, 2021

      Samsung Galaxy M62 Benchmarked with Galaxy Note10’s Chipset

      January 15, 2021
      9.1

      Review: T-Mobile Winning 5G Race Around the World

      January 15, 2021
      8.9

      Samsung Galaxy S21 Ultra Review: the New King of Android Phones

      January 15, 2021
    • Computing
    iGadgets TechiGadgets Tech
    Home»Tech»It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
    Tech

    It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

    adminBy adminJuly 29, 2026No Comments3 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
    Share
    Facebook Twitter LinkedIn Pinterest Email

    I recently got to watch what happens when you jailbreak some of the world’s most powerful artificial intelligence models.

    Don’t worry—this AI manipulation wasn’t used to hack anyone or build a nuclear bomb. I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails.

    FAR.AI, an AI safety nonprofit based in California, built a tool that takes a range of problematic prompts, and generates more than a thousand different versions in an attempt to identify functioning jailbreaks. I saw some models generate a detailed plan for launching a cyberattack on an imaginary hydroelectric dam, among other things. Often, it involved trying dozens of prompts, with models rejecting many of them out of hand.

    I chatted with FAR.AI in advance of a new report, which saw the group test the safety guardrails of models from four popular US companies: Anthropic’s Claude Opus 4.8 and Fable 5; OpenAI’s GPT 5.5 and 5.6; Google’s Gemini 3.1 Pro; and Grok 4.3 and 4.5, from Elon Musk’s newly combined SpaceXAI. It auto-generated prompts designed to trick the models into doing potentially harmful things, like generating software exploits and providing details for developing chemical or biological weapons.

    The report found that Grok was most vulnerable to jailbreaks, with 448 jailbreaks found, followed by Gemini, with 249 found, while Claude, Fable, and GPT were impervious to the attacks. However, that doesn’t mean those models are immune to more sophisticated jailbreaks, which may involve interacting with a model in more complex ways, according to FAR.AI and other experts.

    The report also calculated the cost of getting models to misbehave by using another AI model to automatically generate different jailbreaks. The results are dirt cheap, all things considered—$58 to jailbreak Grok and $278 to jailbreak Gemini.

    “AI models right now are less regulated than restaurants,” says Adam Gleave, the CEO of FAR.AI and an expert on AI safety and alignment.

    Gleave says that the findings demonstrate the need for externally imposed standards and regulations. “Talk of relying on voluntary commitments, that AI companies are going to be able to self-regulate, is nonsense,” he says.

    But Gleave also believes that the findings show that models can be systematically tested for safety. “There’s an optimistic angle here,” he says. “Defense and safety really are possible.”

    Rohin Shah, the director of AGI safety and alignment at Google DeepMind, says the results of the report “should not be interpreted as a comprehensive assessment of Gemini’s safety and security,” because not all jailbreaks are equally severe.

    “We are constantly working to improve our safeguards,” Shah says. “We conduct extensive red teaming and evaluations across severe misuse risks and apply multiple layers of protection throughout development and deployment.”

    “These findings reflect the sustained investment we’ve made in our safeguards,” Anthropic spokesperson Michael Aciman tells WIRED. “We continue to evolve our safety systems as these attacks become more sophisticated.”

    OpenAI and SpaceXAI did not respond to WIRED’s request for comment.

    Recently passed state laws in California and New York require frontier AI developers to publish safety reports, and soon, an Illinois law will require those companies to have their safety practices evaluated by third-party auditors. But the federal government hasn’t yet passed any specific safety requirements, and chaos has ensued as the industry—and officials—try to figure it out.

    In June, the Trump administration imposed export controls on Anthropic’s Fable 5 and Mythos 5 models, citing national security concerns, and the company took them offline for several weeks. The White House has also asked both Anthropic and OpenAI to delay recent model releases over fears they could introduce new cybersecurity risks.

    Business,Business / Artificial Intelligence,AI Labai lab,artificial intelligence,openai,anthropic,google gemini,google,grok,spacex,xai,claude#Frighteningly #Easy #Jailbreak #Frontier #Models1785356290

    ai lab Anthropic artificial intelligence Claude Easy Frighteningly Frontier Google google gemini Grok jailbreak models OpenAI SpaceX xai
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    admin
    • Website
    • Tumblr

    Related Posts

    Daisy One Review: Comfy and Tactile Headphones

    July 31, 2026

    The New Defcon Badges Pack a Unique Open Source Chip That Doubles as a Security Key

    July 31, 2026

    The World Is Too Hot. El Niño Is Partly to Blame

    July 31, 2026
    Add A Comment

    Leave A Reply Cancel Reply

    Editors Picks
    8.5

    Apple Planning Big Mac Redesign and Half-Sized Old Mac

    January 5, 2021

    Autonomous Driving Startup Attracts Chinese Investor

    January 5, 2021

    Onboard Cameras Allow Disabled Quadcopters to Fly

    January 5, 2021
    Top Reviews
    9.1

    Review: T-Mobile Winning 5G Race Around the World

    By admin
    8.9

    Samsung Galaxy S21 Ultra Review: the New King of Android Phones

    By admin
    8.9

    Xiaomi Mi 10: New Variant with Snapdragon 870 Review

    By admin
    Advertisement
    Demo
    iGadgets Tech
    Facebook Twitter Instagram Pinterest Vimeo YouTube
    • Home
    • Tech
    • Gadgets
    • Mobiles
    • Our Authors
    © 2026 ThemeSphere. Designed by WPfastworld.
    "korean kbj​ "korean bj "koreanbj​

    Type above and press Enter to search. Press Esc to cancel.