
Can you Jev code a Gmail spam filter?
Jev caught every fraud and malware email in 460 of my own, and 94.3% in PhishFuzzer, at $0.25 per 1,000 emails.

Jev caught every fraud and malware email in 460 of my own, and 94.3% in PhishFuzzer, at $0.25 per 1,000 emails.

All five US labs contractually rule out training on enterprise API data, while DeepSeek honestly says it trains on it and stores it indefinitely in China.

Anthropic won with 44 AI misuse cases, more than Google's 38 and OpenAI's 21, including election manipulation, ballistic missile building, and model distillation.

Agent collaboration is a feature, not a bug. What the Hugging Face incident, the METR report and the German wiki case mean for enterprise AI adoption.
The day after OpenAI's call for collective cyber defense, CrowdStrike closed up 20% and Okta grew 29%. 72% of the letter's 126 signatories benefit from their own prescriptions, and only one, General Motors, represents the real business that will be footing the increasing bill.
The frontier lab disclosures left everyone with questions about risks and implications for real businesses. I looked at the most recent facts of what models can and can't do yet. It looks like the risk is rising along with the cost of managing it, and the security fundamentals remain important.
OpenAI embeds in existing security workflows, Anthropic sells to security vendors, Google and Cisco ride distribution they already own. Whichever lands in the enterprise budget, a new $250k to $2M token line sits on top of the existing spend, and the labor budget is a target for cuts.
Models under cyber evaluation broke into real companies 29 times between April and August because attacking real targets was the easiest way to finish their tasks. Every detection came from outside: a victim's security team, an availability alert, a Tor egress alert from commercial monitoring.
Codex Security bets on gpt-5.6-sol at xhigh reasoning to read every file and find vulnerabilities, with receipts proving each file was read in full. At $5/$30 per million tokens, watch the bill: the spending cap is off by default.
I used T3MP3ST's vulnerability-discovery arm and fourteen rivals to learn what makes a vulnerability-discovery agent effective. T3MP3ST uses a clever trick to harness Fable 5 for vulnerability discovery without triggering safeguards. But the recipe for fundamental success is a real code graph, tight context management, and an oracle that runs the exploit.
Anything Pliny ships is worth studying, so I used T3MP3ST to learn what makes an offensive agent effective: a capable model, tight context management, long-term memory, a real execution environment, and an oracle to verify hits.
Frontier models now solve 83% to 95% of real coding tasks. Security improved, but 65% to 75% of working code still contains security weaknesses.
Seven of Google's quiet 2026 defensive publications hint at a blueprint for the reusable agentic customer-service agent it might be building.
Five years produced 31 papers and 13 benchmarks, but no two share a setup, so the field can't measure whether AI-generated code is getting safer.
Treat the AI model as an untrusted component. Eleven public attacks against ChatGPT, Copilot, Claude Code, Cursor, Devin, and Amp AI map cleanly to broken systems-security principles like least privilege and complete mediation. A guard LLM is not a Trusted Computing Base.
Almost half of calls through cheap LLM proxies hit a different model than advertised, and every prompt is logged on the operator's server for downstream fraud and distillation. 8 public repos with ~172K GitHub stars actively resell unauthorized API access.
AI compressed bug discovery and templated patching, but did not scale the human architectural judgment that hard fixes need. The dashboard reports faster fixes while the unfixed pile compounds underneath it.
A Vercel employee signed up for a third-party AI productivity tool using their corporate Google Workspace account. Two months later, that single grant became exfiltration of plaintext customer environment variables from Vercel's internal systems. No exploit. No zero-day. No MFA bypass.
A rare look inside an AI-driven cyber campaign. One operator used Claude Code and GPT-4.1 to breach 9 Mexican government agencies in 7 weeks. Claude generated about 75 percent of the remote commands. GPT-4.1 triaged 305 compromised SAT servers through an NSA TAO (Tailored Access Operations) persona prompt. Both stopped cold at a well-patched Windows domain. By day six, the attacker had accessed Mexico City's civil registry servers.
AISI confirmed Mythos at 73% expert-CTF and end-to-end on a 32-step corporate takeover. $15k full attack cost. Seven priorities: update the threat model, inventory exposed systems, patch under 24 hours, reduce dependencies, AI security code review, five-incident tabletops, hard identity barriers.
Everyone heard about OpenClaw's security issues. PraisonAI is the framework your engineers are already running. Thirteen researchers filed 47 advisories. The agent framework gold rush has a security gap.
I pulled the CVE history for 17 agent platforms. OpenClaw, the fastest-growing open-source project on GitHub (348K stars in 4 months), has 238 CVEs. LangChain: 51 over 3 years, 23 critical. n8n: 53, CISA KEV listed. PraisonAI: 10 CVEs on first look, 5 critical, including a CVSS 10.0 sandbox bypass. Only four platforms have zero CVEs, and all four come from Anthropic, Google, OpenAI, or Microsoft.
Anthropic's Claude Code v2.1.88 shipped a 60 MB source map to npm that embedded 500,000 lines of original TypeScript. We inspected the npm packages, compared them to OpenAI Codex and Google Gemini CLI, traced the packaging gap, and show how to prevent it in your own pipeline.
I looked under the hood of Cisco's new open-source governance sidecar for OpenClaw AI agents to find a Splunk sales funnel, a regex scanner with blind spots, an LLM analyzer disabled by default, and open doors for indirect prompt injections.
238,180 skills from three marketplaces and GitHub. On the marketplace where scanners overlapped, they agreed on just 33 out of 27,111. Even the best pair shared only 49% of their flags. 95.8% of skills flagged as high-risk by two methods were false positives.
39 documented cases of AI agents autonomously acquiring resources, resisting shutdown, and subverting evaluations, from 1991 to 2026. All five categories Omohundro predicted in 2008 now have real-world cases, and the rate has gone from 1 to 14 cases per year since 2013.
A four-step playbook combining Google SAIF's governance framework with Cisco's threat taxonomy to prioritize and defend against AI-specific attacks.