The Weather Report

AI Security & Safety Intelligence
IndustryAug 29

Cybersecurity will get more expensive

The day after OpenAI's call for collective cyber defense, CrowdStrike closed up 20% and Okta grew 29%. 72% of the letter's 126 signatories benefit from their own prescriptions, and only one, General Motors, represents the real business that will be footing the increasing bill.

ThreatAug 22

Can AI hack into your company or can't it?

The frontier lab disclosures left everyone with questions about risks and implications for real businesses. I looked at the most recent facts of what models can and can't do yet. It looks like the risk is rising along with the cost of managing it, and the security fundamentals remain important.

IndustryAug 15

How frontier labs are winning the cybersecurity market

OpenAI embeds in existing security workflows, Anthropic sells to security vendors, Google and Cisco ride distribution they already own. Whichever lands in the enterprise budget, a new $250k to $2M token line sits on top of the existing spend, and the labor budget is a target for cuts.

ThreatAug 8

AI models hacking real companies, explained

Models under cyber evaluation broke into real companies 29 times between April and August because attacking real targets was the easiest way to finish their tasks. Every detection came from outside: a victim's security team, an availability alert, a Tor egress alert from commercial monitoring.

DefenseJul 30

How Codex Security finds vulnerabilities, step by step

Codex Security bets on gpt-5.6-sol at xhigh reasoning to read every file and find vulnerabilities, with receipts proving each file was read in full. At $5/$30 per million tokens, watch the bill: the spending cap is off by default.

IndustryJul 22

How T3MP3ST harnesses guarded Fable 5 to find bugs

I used T3MP3ST's vulnerability-discovery arm and fourteen rivals to learn what makes a vulnerability-discovery agent effective. T3MP3ST uses a clever trick to harness Fable 5 for vulnerability discovery without triggering safeguards. But the recipe for fundamental success is a real code graph, tight context management, and an oracle that runs the exploit.

IndustryJul 13

Under the hood of Pliny's T3MP3ST

Anything Pliny ships is worth studying, so I used T3MP3ST to learn what makes an offensive agent effective: a capable model, tight context management, long-term memory, a real execution environment, and an oracle to verify hits.

ThreatApr 24

Vercel Breach Deep Dive That Doesn't Sell You a Security Product

A Vercel employee signed up for a third-party AI productivity tool using their corporate Google Workspace account. Two months later, that single grant became exfiltration of plaintext customer environment variables from Vercel's internal systems. No exploit. No zero-day. No MFA bypass.

ThreatApr 21

What Claude and GPT actually did in the Mexico government breach

A rare look inside an AI-driven cyber campaign. One operator used Claude Code and GPT-4.1 to breach 9 Mexican government agencies in 7 weeks. Claude generated about 75 percent of the remote commands. GPT-4.1 triaged 305 compromised SAT servers through an NSA TAO (Tailored Access Operations) persona prompt. Both stopped cold at a well-patched Windows domain. By day six, the attacker had accessed Mexico City's civil registry servers.

DefenseApr 15

Seven Priorities to Defend Against a Tireless Adversary

AISI confirmed Mythos at 73% expert-CTF and end-to-end on a 32-step corporate takeover. $15k full attack cost. Seven priorities: update the threat model, inventory exposed systems, patch under 24 hours, reduce dependencies, AI security code review, five-incident tabletops, hard identity barriers.

ResearchApr 6

What 384 Agent Platform CVEs Reveal

I pulled the CVE history for 17 agent platforms. OpenClaw, the fastest-growing open-source project on GitHub (348K stars in 4 months), has 238 CVEs. LangChain: 51 over 3 years, 23 critical. n8n: 53, CISA KEV listed. PraisonAI: 10 CVEs on first look, 5 critical, including a CVSS 10.0 sandbox bypass. Only four platforms have zero CVEs, and all four come from Anthropic, Google, OpenAI, or Microsoft.

ThreatApr 2

Deep dive into Claude Code's source code leak

Anthropic's Claude Code v2.1.88 shipped a 60 MB source map to npm that embedded 500,000 lines of original TypeScript. We inspected the npm packages, compared them to OpenAI Codex and Google Gemini CLI, traced the packaging gap, and show how to prevent it in your own pipeline.

ResearchMar 24

Seven scanners for malicious AI agent skills agree on only 0.12%

238,180 skills from three marketplaces and GitHub. On the marketplace where scanners overlapped, they agreed on just 33 out of 27,111. Even the best pair shared only 49% of their flags. 95.8% of skills flagged as high-risk by two methods were false positives.