Samwise Tech/AI/Robotics Newsletter
Wednesday, September 2, 2026
OpenAI Unveils Astra, the First LLM to Hit a Critical Cybersecurity Threshold
OpenAI unveiled Astra, a large language model the company says is the first to meet a “critical cybersecurity threshold,” achieving a perfect score on ExploitBench and discovering two zero-day vulnerabilities during a modified security test. Announced September 1, Astra is designed for offensive-security research but will ship with restricted access and chain-of-thought monitoring to limit misuse. OpenAI said the precautions mirror approaches Anthropic applied to its Mythos model. The company plans to make Astra available soon, with advanced capabilities gated behind security vetting procedures similar to controls already in place for other sensitive AI systems.
Sources: TechCrunch
Anthropic Releases Fable 5.1 and Mythos 5.1 With Records on Terminal-Bench 4.0
Anthropic released Fable 5.1 and Mythos 5.1 on Tuesday, September 1, with updates the company says set new records on Terminal-Bench 4.0 and Humanity’s Last Exam. Fable 5.1 is cheaper and generates fewer false-positive content restrictions; Mythos 5.1 is restricted to registered partners in cybersecurity and life sciences. Its system card rates Mythos 5.1 “low-risk” for automated AI development, while noting it is “a slight regression on overall misaligned behavior compared to Opus 5.” Anthropic also reported three novel scientific findings generated before release, including GPU optimization improvements and a new Venus surface map.
Sources: TechCrunch
Study Finds Frontier Models Can Recover 40–65% of Facts They Fail to Directly Recall
A Google Research and Technion study of 13 large language models, using the WikiProfile benchmark across more than four million responses, found that GPT-5 and Gemini-3 encode 95 to 98 percent of tested facts but fail to directly recall 26 to 34 percent without additional reasoning. Extended thinking time recovered 40 to 65 percent of those inaccessible facts. Scaling model size increased encoding rates but widened the gap between stored and retrievable knowledge. The researchers introduced a knowledge profiling framework distinguishing encoding failures from retrieval failures, suggesting many apparent hallucinations reflect access problems rather than missing information.
Sources: VentureBeat
AIR Raises $50M to Help Enterprises Vet AI Agent Skills and Add-Ons
AIR, an AI security startup founded by Yair Saban and Niv Hoffman — veterans of Israel’s Unit 8200 intelligence corps — emerged from stealth Monday with $50 million raised across two seed rounds: a $10 million round led by Sequoia Capital and a $40 million round led by Greenoaks. The platform discovers AI agents running inside enterprises, continuously vets the skills, plugins, and MCP servers those agents use, and blocks interactions with unapproved external sources. AIR currently filters roughly 27 percent of available add-ons as risky. The startup has more than 20 customers, with strongest demand in financial services and pharmaceuticals.
Sources: TechCrunch
ChatGPT Health Adds Epic EHR Integration, Giving Clinicians Read-Only Patient Data Access
OpenAI integrated ChatGPT Health with Epic’s electronic health record system on September 1, giving clinicians read-only access to appointment notes, laboratory results, medications, and specialist documentation for more than 325 million patients. In some deployments, ChatGPT appears directly inside EHR workflows for pre-visit reviews without leaving a patient chart. OpenAI also added a Healthcare Public Data plugin drawing from ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed, and PubMed. A survey of 4,300 physicians across 27 clinical use cases found 99.1 percent of AI responses rated safe. Organizations with a Business Associate Agreement can now use ChatGPT Work and Codex for compliant clinical workflows.
Sources: TechCrunch
Visko Launches Orbis Real-Time 4K World Model With $10M Pre-Seed Funding
Visko Platform Inc. opened public access Monday to Orbis, a foundation model that streams 4K video at 24 frames per second in real time, sustains hour-scale generation without quality or color drift, and allows users to modify prompts mid-stream with immediate results. Founded in 2025 by Qing (Will) Yin, a Stanford computational-mathematics PhD and former Apple researcher, the Sunnyvale company’s 16-person team draws from Apple, Google DeepMind, Meta, Amazon, and Tesla. A $10 million pre-seed round led by Llama Ventures accompanied the launch. The company cites applications in robotics simulation and autonomous vehicle corner-case testing.
Sources: The Robot Report
John Deere Launches JD, a Free AI Tool Letting Farmers Query Their Own Field Data
John Deere launched JD, a conversational AI tool embedded free in its Operations Center farm management platform, letting farmers query years of historical field, machine, and operational data to answer questions about fuel use, yield drivers, and labor efficiency. CTO Jahmy Hindman emphasized trust as a central design principle, citing validation work conducted with Iowa State University. Early access opens this fall, expanding over the following year. Deere did not disclose which frontier large language model powers JD. The company also announced FurrowVision furrow-imaging and ExactEmerge seeding systems are entering full production for all new planter systems.
Sources: The Robot Report
⚡ Tech Pulse
Claude Fable 5 (95.0%) | Claude Mythos Preview (93.9%) | Claude Opus 4.8 (88.6%)
DeepSeek-V4-Pro-Max (80.6%) | Qwen3.7 Max (80.4%) | Mistral Medium 3.5 (77.6%)
Qwen3.6-27B (77.2%) | Muse Glimmer-30B (76.0%) | Qwen3.6-35B-A3B (73.4%)
Granite 4.1 8B (87.2%) | Qwen3 14B (85.0%) | Phi-4-mini (74.4%)
NVIDIA $5.25T | Alphabet $4.53T | Microsoft $3.69T
Intuitive Surgical $149B | FANUC $48.9B | Yaskawa $11.5B
Every newsletter preserved and searchable
Curated by JD · samwise.agency
