Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Nscale buys Anyscale as it seeks to own more of the AI compute stack

    July 31, 2026

    Daisy One Review: Comfy and Tactile Headphones

    July 31, 2026

    How the 60 Minutes Editorial Process Is Supposed to Work

    July 31, 2026
    Facebook Twitter Instagram
    • Tech
    • Gadgets
    • Spotlight
    • Gaming
    Facebook Twitter Instagram
    iGadgets TechiGadgets Tech
    Subscribe
    • Home
    • Gadgets
    • Insights
    • Apps

      Nscale buys Anyscale as it seeks to own more of the AI compute stack

      July 31, 2026

      Netflix lands global streaming deal for ‘The Walking Dead’

      July 31, 2026

      Meta says AI is making it easier to build new apps — and more are coming

      July 31, 2026

      When will fusion power startup Commonwealth Fusion Systems go public?

      July 31, 2026

      Spotify launches ‘User Notes’ to let users add memories to songs

      July 31, 2026
    • Gear
    • Mobiles
      1. Tech
      2. Gadgets
      3. Insights
      4. View All

      Daisy One Review: Comfy and Tactile Headphones

      July 31, 2026

      The New Defcon Badges Pack a Unique Open Source Chip That Doubles as a Security Key

      July 31, 2026

      The World Is Too Hot. El Niño Is Partly to Blame

      July 31, 2026

      SelectBlinds Promo Codes & Coupons: Save on Custom Window Treatments

      July 31, 2026

      March Update May Have Weakened The Haptics For Pixel 6 Users

      April 2, 2022

      Project 'Diamond' Is The Galaxy S23, Not A Rollable Smartphone

      April 2, 2022

      The At A Glance Widget Is More Useful After March Update

      April 2, 2022

      Pre-Order The OnePlus 10 Pro For Just $1 In The US

      April 2, 2022

      Motorola Edge+ Review: It Checks A Lot Of Boxes

      April 2, 2022

      This Smartphone Concept Design Is Different… In A Good Way

      April 2, 2022

      Twitter Just Made Searching Your Direct Messages Better

      April 2, 2022

      That Netflix Price Hike Is Starting To Take Place

      April 2, 2022

      Latest Huawei Mobiles P50 and P50 Pro Feature Kirin Chips

      January 15, 2021

      Samsung Galaxy M62 Benchmarked with Galaxy Note10’s Chipset

      January 15, 2021
      9.1

      Review: T-Mobile Winning 5G Race Around the World

      January 15, 2021
      8.9

      Samsung Galaxy S21 Ultra Review: the New King of Android Phones

      January 15, 2021
    • Computing
    iGadgets TechiGadgets Tech
    Home»Tech»You Can Now Sound the Alarm on AI Behaving Badly
    Tech

    You Can Now Sound the Alarm on AI Behaving Badly

    adminBy adminJuly 1, 2026No Comments3 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    You Can Now Sound the Alarm on AI Behaving Badly
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Writing AI Lab each week means I occasionally encounter AI models that behave badly and bizarrely. Usually, there’s nothing to be done about it, save for sharing those tales with you. But that could soon change.

    A group of AI researchers has set up a crowdsourced website, Flaw Reporting for AI (FLARE-AI), for reporting and tracking AI harms. If, for example, a chatbot generates malware or a bomb-making recipe, leaks personal information, or triggers delusional thinking in users, FLARE-AI could be used to sound the alarm. The open source code behind the system allows others to verify an issue and route reports to model makers, as well as organizations like MITRE, a nonprofit that tracks problems with technical systems. It’s a bit like Downdetector, which compiles real-time user reports for global service outages affecting things like apps and websites.

    The website is another step in the group’s ongoing work with AI reporting, which I first wrote about last year. Members of the group also consulted on a congressional bill announced in June, which would see the US government take a central role in tracking this kind of AI misbehavior.

    “Right now, there is no centralized, accountable way to report flaws in AI systems,” says Avijit Ghosh, an artificial intelligence policy researcher at HuggingFace who co-led development of FLARE-AI with computer scientists Elaine Zhu and Shayne Longpre.

    The alarm system was developed in collaboration with 49 AI experts from 32 different organizations. In a paper outlining the work, the researchers argue that their initiative could prove crucial as AI is adopted more widely and as agentic systems gain greater power. The lack of a consistent way to report AI flaws is a significant problem, they believe.

    “I think it’s a really good initiative,” says Jessica Ji, a researcher at the think tank Center for Security and Emerging Technology. Ji says the researchers are right to note that existing reporting mechanisms are fragmented and that AI models are black boxes. “I’m in support of anything that makes AI more transparent,” she says.

    Though bugs and cybersecurity problems get a lot of attention—especially of late—Ghosh tells me that problems with AI systems span topics like psychological harm, discrimination or bias, and misinformation. He adds that different companies have different standards around such issues, which means some problems go unrecognized. “In the absence of a coordinated disclosure system, there are no external mechanisms to enforce transparency,” Ghosh says.

    A spate of recent incidents involving popular AI tools shows how easily the technology can go bad.

    This week, a company called LayerX disclosed a way to dupe AI-infused web browsers, including OpenAI’s Atlas and Perplexity’s Comet, into vaulting their guardrails. Convincing the AI model behind the browser that it was playing a game, for example, could lead to the browser going rogue and trying to hack a website. (The companies responsible for the affected browsers have fixed the issue, LayerX says.) And this April, Johann Rehberger, a security researcher, discovered a way to trick Claude into divulging personal data using images generated by ChatGTP.

    AI introduces bizarre new kinds of problems, too. Last year, OpenAI was forced to update its models after it discovered that they were overly sycophantic, which sometimes appeared to encourage delusional thinking.

    Rumman Chowdhury, the CEO and founder of Humane Intelligence PBC, says FLARE-AI could be a useful way for many AI developers to implement ways of reporting issues with their tools. But she adds that such initiatives often come with serious challenges.

    Business,Business / Artificial Intelligence,AI Labai lab,artificial intelligence,cybersecurity,congress,chatbots,openai,claude,chatgpt,anthropic#Sound #Alarm #Behaving #Badly1782930197

    ai lab alarm Anthropic artificial intelligence Badly Behaving chatbots ChatGPT Claude congress cybersecurity OpenAI Sound
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    admin
    • Website
    • Tumblr

    Related Posts

    Daisy One Review: Comfy and Tactile Headphones

    July 31, 2026

    The New Defcon Badges Pack a Unique Open Source Chip That Doubles as a Security Key

    July 31, 2026

    The World Is Too Hot. El Niño Is Partly to Blame

    July 31, 2026
    Add A Comment

    Leave A Reply Cancel Reply

    Editors Picks
    8.5

    Apple Planning Big Mac Redesign and Half-Sized Old Mac

    January 5, 2021

    Autonomous Driving Startup Attracts Chinese Investor

    January 5, 2021

    Onboard Cameras Allow Disabled Quadcopters to Fly

    January 5, 2021
    Top Reviews
    9.1

    Review: T-Mobile Winning 5G Race Around the World

    By admin
    8.9

    Samsung Galaxy S21 Ultra Review: the New King of Android Phones

    By admin
    8.9

    Xiaomi Mi 10: New Variant with Snapdragon 870 Review

    By admin
    Advertisement
    Demo
    iGadgets Tech
    Facebook Twitter Instagram Pinterest Vimeo YouTube
    • Home
    • Tech
    • Gadgets
    • Mobiles
    • Our Authors
    © 2026 ThemeSphere. Designed by WPfastworld.
    "korean kbj​ "korean bj "koreanbj​

    Type above and press Enter to search. Press Esc to cancel.