Article

    Cyber News / Article / “Sorry, I can’t help with that”: How your guardrails might become the attacker’s best friend

    “Sorry, I can’t help with that”: How your guardrails might become the attacker’s best friend
    Da
    David J. Bianco-15 days ago

    “Sorry, I can’t help with that”: How your guardrails might become the attacker’s best friend

    Welcome to this week’s edition of the Threat Source newsletter.

    Hello, everyone. Long time reader, first time writer here at the Threat Source newsletter! I wanted to start out by introducing myself. My colleague and friend Mick Baccio set the bar pretty highlast week, so I was planning to tell you all about myself, including:

    Unfortunately, myeditorsays we don’t have the “space” for that, the MIT thing might open me up to “liability,” and it’s not the kind of “professional image” we strive for here at Talos. (I'm watching. Always watching. -Amy)

    So instead, I’ll just play it safe and say that I’ve been in the security field for a little over 30 years now, mostly concentrating on the defensive side (Go, Team Blue!). I’ve helped set up SOCs, run threat hunting teams, and even published afew thingsyou might haveheard of.

    Speaking of things I’ve published, I’ve written before about theAttacker’s Dilemma. The idea that defenders have inherent advantages over attackers runs contrary to what most of us have heard throughout our careers. An attacker must evade monitoring and technical controls at every step of their attack lifecycle, because the defender only needs to noticeoncein order to respond and prevent them from achieving their goal. This is one of the most important advantages of any security team has, but we are currently witnessing a self-imposed erosion of this advantage through the rise of poorly-designed AI guardrails.

    I’m not opposed to guardrails, but we have to carefully consider what we’re guarding against and where we deploy them. As I explored in a recent piece onThe Safety Penalty, by allowing third-party AI providers to implement and control safety filters and the policies behind them, we may in fact be helping the attacker. If agentic SOC process experience refusals, it can slow or even halt investigations. Of course, these should get flagged for human intervention, but that takes time and may give the attacker breathing room in which to complete their mission.

    It may turn out that thewhereof the guardrails is even more important than thewhat. Operational sovereignty relies on having control of our own limits. Any vision of an agentic SOC must allow the security teams to customize the guardrails according to their own threat model. They should also have the flexibility to temporarily remove specific safeguards under authorized circumstances, something you won’t get with guardrails from a frontier provider. These controls belong inside your organization’s agentic harness where you can set the policies and technical controls to allow you to analyze threats while ensuring your agents stay within their lanes.

    Ultimately, operational sovereignty means engaging with the reality of the threat landscape, ensuring that the adversary can’t derail the defender’s investigation and response processes, either accidentally or intentionally. We need to move toward a model where each organization can choose the guardrails that work for them, rather than having inflexible guardrails chosen for them.

    Cisco Talos recentlyevaluated 66 large language model (LLM) and reasoning combinationsto see if we could find a clear winner for security operations. Instead, we found that selecting the right model is a complex balancing act between efficacy, speed, cost, and consistency. Cranking up a model's reasoning effort doesn't guarantee better analysis and can actually degrade performance. Ultimately, we developed a repeatable methodology to help organizations navigate these tradeoffs for their own workflows.

    Choosing an AI model based solely on generic leaderboard scores is a recipe for operational disaster. An exceptionally smart model might cost a fortune, take half an hour to analyze a single log, or completely fail to format its output. Assuming more compute power equals better results is a costly trap, as higher reasoning settings sometimes produce weaker or blocked responses. Defenders must remember that prompts, analyst personas, and model consistency drastically alter an investigation's outcome.

    Test models against your organization’s specific workflows before deploying them. Build a focused set of representative cases and test them multiple times using the exact prompts and tools your analysts will actually use. Track the quality, cost, time, consistency, and usable-answer rates in a simple spreadsheet to expose the real-world tradeoffs. Finally, establish acceptable thresholds for these variables to eliminate underperforming models, and regularly revisit your decisions as AI technology and pricing inevitably shift.

    ToxicPanda banking trojan matures into enterprise threatToxicPanda 2.0 expands substantially on its predecessor, adding 167 remote commands and broadening its targeting from 16 financial institutions to 349 banking, e-wallet, and cryptocurrency applications. (Dark Reading)

    Interpol's Jackal IV disrupts West African crime infrastructureLaw enforcement from 22 countries across six continents worked together to arrest 58 suspects and identify 263 more. The first two Jackal operations in 2022 and 2023 led to approximately 200 arrests in total and millions of dollars more in seized assets. (Dark Reading)

    First malware built specifically for car head units fuels botnetResearchers have found what appears to be the first malware specifically designed for car head units, with links to the notorious BadBox botnet, on an Android-powered aftermarket infotainment system made by Chinese company DoFun, which is widely used in China and other APAC countries. (SecurityWeek)

    A Tale of Two SOCs: Insights From Two Red Team AssessmentsA CISA red team fully compromised two critical infrastructure organizations at the domain level and reached sensitive business systems and cloud resources. Organization A failed to detect or contain the activity. Organization B rapidly identified initial compromise attempts, isolated affected systems, and forced the red team into an assume breach model. (CISA)

    NovaCookies campaigns abuse genuine Docusign notifications to steal M365 sessionsThe $320/month service is a subscription-based phishing platform that facilitates real-time M365 session theft. The kit has been used to target hundreds of organizations across multiple sectors in the U.S., the U.K., Canada, Germany, and more. (The Hacker News)

    JavaScript obfuscation: From party trick to phishing kitWe've spent a lot of time pulling apart suspicious JavaScript from phishing kits, malware packages, compromised sites, and more. Learn the basics of what obfuscation is, why a researcher would try to reverse it, and several ways to approach the problem.

    The safety penalty: Reclaiming operational sovereignty in the age of AIAs frontier AI models become increasingly restrictive, security teams are facing a "safety penalty" that hampers real-time incident response. Discover how organizations can move toward operational sovereignty to ensure their defensive AI keeps pace with unconstrained adversaries.

    Back-to-school cybersecurity: Protecting education networks from ransomware and threatsAs the new academic year begins, school districts face a surge in cybersecurity threats, from phishing attacks and ransomware to student experimentation with network devices. In this episode, Amy sits down with Cisco Talos expert Pierre Cadieux to discuss practical strategies for IT practitioners.

    SHA256: 9f1f11a708d393e0a4109ae189bc64f1f3e312653dcf317a2bd406f18ffcc507MD5: 2915b3f8b703eb744fc54c81f4a9c67fTalos Rep:https://talosintelligence.com/talos_file_reputation?s=9f1f11a708d393e0a4109ae189bc64f1f3e312653dcf317a2bd406f18ffcc507Example Filename: VID001.exeDetection Name: W32.9F1F11A708-100.SBX.TG**

    SHA256: e7e784cae8d37f12a5af0bc9b3975c8d3e668142e9c6b0b365ed4f4e80933c47MD5: a4480423617d0b0d3b38c8471cbf594cTalos Rep:https://talosintelligence.com/talos_file_reputation?s=e7e784cae8d37f12a5af0bc9b3975c8d3e668142e9c6b0b365ed4f4e80933c47Example Filename: client32.exeDetection Name: W32.Trojan.29ev.1201

    SHA256: c4dd71e347a076ba24bdd2d0ee532ef991c1ef25a2431a19f850942ba2ab16b2MD5: 9a47c4d379998ade2f8f99e23a630c06Talos Rep:https://talosintelligence.com/talos_file_reputation?s=c4dd71e347a076ba24bdd2d0ee532ef991c1ef25a2431a19f850942ba2ab16b2Example Filename: WCInstaller_NonAdmin.exeDetection Name: W32.C4DD71E347-95.SBX.TG

    SHA256: 38d053135ddceaef0abb8296f3b0bf6114b25e10e6fa1bb8050aeecec4ba8f55MD5: 41444d7018601b599beac0c60ed1bf83Talos Rep:https://talosintelligence.com/talos_file_reputation?s=38d053135ddceaef0abb8296f3b0bf6114b25e10e6fa1bb8050aeecec4ba8f55Example Filename: content.jsDetection Name: W32.38D053135D-95.SBX.TG

    SHA256: 9896a6fcb9bb5ac1ec5297b4a65be3f647589adf7c37b45f3f7466decd6a4a7fMD5: 38de5b216c33833af710e88f7f64fc98Talos Rep:https://talosintelligence.com/talos_file_reputation?s=9896a6fcb9bb5ac1ec5297b4a65be3f647589adf7c37b45f3f7466decd6a4a7fExample Filename: SECOH-QAD.exeDetection Name: Win.Tool.Procpatcher::1201

    SHA256: a31f222fc283227f5e7988d1ad9c0aecd66d58bb7b4d8518ae23e110308dbf91MD5: 7bdbd180c081fa63ca94f9c22c457376Talos Rep:https://talosintelligence.com/talos_file_reputation?s=a31f222fc283227f5e7988d1ad9c0aecd66d58bb7b4d8518ae23e110308dbf91Example Filename:d4aa3e7010220ad1b458fac17039c274_62_Exe.exeDetection Name: Win.Dropper.Miner::95.sbx.tg**

    From engaging with cybercriminals to surviving a live Flamin’ Hot Cheetos taste test, Hazel reflects on the latest Beers with Talos with Azim, where they cover the full spectrum of what it takes to gather threat intel.

    In this week's newsletter, new author Mick Baccio introduces himself and explores the operational and security implications of the new White House memorandum regarding private sector participation in government-authorized offensive cyber operations.

    In this edition of the Threat Source newsletter, William reflects on the “Make Hazel a Hacker” segment in Beers with Talos, and how cybersecurity is a field where questions can lead to multiple correct answers.

    Original source