Wednesday, July 29, 2026 · 10 curated articles

Editor's Picks
The loudest story in AI today is not another trillion-parameter release. It is OpenAI Codex lead Tibo (@thsottiaux) hitting the usage-reset button again—for every ChatGPT Work and Codex subscriber—then explaining why GPT-5.6 Sol drained limits faster than people expected. Sol runs longer, makes extra tool calls, and coordinates messy multi-tool and subagent workflows. At the same High effort setting it can spend more tokens than GPT-5.5 did. Programmatic tool calling (code mode) lets it work while tools are pending, which also means more responses per turn and more cached input. Median users were fine; power users on hard tasks got emptied. The team says typical Sol use should last about 18% longer now, and the five-hour limit paused during the investigation is coming back tomorrow—replies immediately begged them not to.
What makes this more than a status post is the culture that formed around the button. People call him Lord Tibo. Fans built Codex Radar and Tibo Reset Radar apps that watch his feed and try to forecast the next gift. Sam Altman once joked that if a tweet got one like, Tibo would reset limits. Tibo himself posted: “One day we created the reset button and the rest is history.” Claude users still rage about rate limits without a comparable folk hero who can top everyone up with a tweet. On the Codex side, a product ops lever became a public ritual: celebration, memes, pre-reset sprint mode, and a shared identity signal—do you refresh Tibo’s timeline when your bar hits red?
Under the jokes sits a structural fact: agent capability and token efficiency are not moving in lockstep. The same behaviors that make Sol good at hard work also make bills explode. Launch cost models that optimize for median chat miss the long tail of overnight autonomous coding. Elsewhere today, Kimi K3 drops 2.8T open weights, Claude Mythos halves HAWK’s effective key strength in 60 hours, and audits show about 12% of questions on GPQA-class benchmarks were broken. Closed-source agent life is now half engineering, half usage politics and meme governance; open weights and cleaner evals are the other half of the pressure. The useful question for builders is not “when does Tibo press it next?” but whether your stack measures real agent runs—Prefactor-style eval layers, hard tasks, verifiable crypto and benchmark hygiene—so capability jumps turn into systems you can still afford and trust.
Developer Tools
Once coding agents are daily infrastructure, quotas, resets, and supply-chain hardening are product features, not back-office noise. Today’s tools span both “how do we stay within budget?” and “how do we stay secure while agents touch the repo?”
Tibo Hits Reset Again: How a Quota Button Became AI Coding Subculture
I've reset usage limits for all ChatGPT Work and Codex users.
One day we created the reset button and the rest is history.
OpenAI’s Tibo Sottiaux reset usage for all ChatGPT Work and Codex users and published a plain-spoken autopsy of GPT-5.6 Sol’s burn rate: longer runs, more tool calls, heavier High-effort token use than GPT-5.5, and code mode that multiplies responses and cached input while tools and web searches run. Plans were not cut; typical Sol endurance should improve by roughly 18%, with the temporary five-hour limit returning after the investigation. Community reaction already looks like folk religion—Lord Tibo memes, third-party “reset radars,” and pre-reset Ultra/fast sprints—because the button is predictable enough to plan around. That is the new usage economics of agents: median metrics can look healthy while the long tail of hard work empties the meter, and a human with a dashboard becomes both capacity manager and community mascot.
Source: X / Tibo @thsottiaux

GitHub Strengthens npm and Actions Security to Thwart Supply Chain Attacks
changes we've shipped across npm and GitHub Actions over the past few months to disrupt supply chain attack techniques
disrupt supply chain attack techniques and limit their impact
GitHub has implemented a series of updates across the npm registry and GitHub Actions over the past few months to proactively disrupt supply chain attack techniques. These security enhancements focus on mitigating the impact of malicious actors attempting to compromise the software development lifecycle through automated workflows. By strengthening the integration between package management and GitHub Actions, the platform aims to provide a more resilient environment for global developers. The implemented changes specifically target common vulnerabilities that attackers exploit to inject malicious code into widely used open-source projects. This ongoing security initiative underscores the critical necessity of maintaining data integrity and trust within the modern software ecosystem. Developers are urged to leverage these improved tools to safeguard their codebases against increasingly sophisticated and automated threats.
Source: The GitHub Blog

Foundation Models
Foundation models still trade scale and architecture for a higher ceiling. Open-weight Kimi K3 pushes trillion-class MoE and million-token context into downloadable territory, while closed coding agents force a separate conversation about burn rate and efficiency.
Kimi K3: A 2.8T Parameter MoE Model with 1-Million-Token Context Window
Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window.
Stable LatentMoE, which effectively activates 16 of 896 routed experts per token... yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2.
Kimi K3 represents a significant leap in open-source AI, featuring a 2.8-trillion parameter Mixture-of-Experts architecture with 104 billion activated parameters per token. The model integrates Kimi Delta Attention and Attention Residuals to optimize information flow across its massive 1-million-token context window and deep architecture. Utilizing Stable LatentMoE to activate 16 out of 896 experts, the system achieves a 2.5x improvement in overall scaling efficiency compared to its predecessor, Kimi K2. Post-training focuses on reinforcement learning across coding, agentic, and reasoning domains, enabling the model to handle long-horizon tasks and complex compositional generalization. While currently trailing specific proprietary models like Claude Fable 5 and GPT-5.6 Sol, Kimi K3 sets a new benchmark for open frontier intelligence. The full model weights have been released to the research community to accelerate the global adoption of high-performance foundation models and vision-integrated intelligence.
Source: HuggingFace Papers

Research
This section covers hard technical work: Claude Mythos finding new crypto weaknesses, HiFi-UMI training robot policies from high-fidelity human data, and audits that remove broken questions from GPQA-class benchmarks—raising the ceiling and cleaning the scoreboard at once.
Claude Mythos Preview Identifies New Weaknesses in Cryptographic Algorithms
Mythos was able to improve the best-known attack on it in just 60 hours of work—effectively cutting its key strength in half.
Mythos found a way to break one such weaker version, and eliminated one of the guesses an attacker needs to make, improving the speed of the previous best attacks by 200-800×.
Claude Mythos Preview has successfully identified mathematical flaws in cryptographic algorithms, including a significant attack that reduces the security of the post-quantum signature scheme HAWK. This discovery effectively cut the key strength of HAWK in half after just 60 hours of autonomous work, despite the algorithm previously surviving two years of expert human review. Additionally, the model improved attacks on a reduced version of the Advanced Encryption Standard (AES), increasing the speed of existing attacks by 200 to 800 times. While these research findings represent substantial advances in AI-driven cryptanalysis, they do not currently affect any production systems or widely deployed software. The research highlights the growing capability of foundation models to find flaws in the fundamental building blocks of digital security beyond simple implementation errors. These results suggest that AI will play a critical role in evaluating and strengthening future cryptographic standards.
Source: Hacker News

HiFi-UMI: High-Fidelity Data for Deployable Robot Manipulation Policies
It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure.
A policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation.
HiFi-UMI achieves 3 mm workspace-local end-effector accuracy without external tracking infrastructure, enabling the training of deployable robot policies without real-robot data anchors. The system utilizes head-mounted offline stereo-inertial SLAM, native relative pose, and microsecond-synchronized GPIO triggers to capture high-fidelity human demonstrations. Policies post-trained solely on this robot-free data match in-domain teleoperation performance across three major backbones, including StarVLA-QwenPI and OpenPI-pi_0.5, with success-rate differences ranging from -2.5 to +3.1 percentage points. The researchers have open-sourced HiFi-UMI-2K, a dataset containing 2,000 hours of microsecond-synchronized demonstrations validated through simulation replay. Pre-training on 4,000 hours of this corpus reduced action error by 41% on unseen tasks and increased real-robot success rates by 18.1 percentage points on the StarVLA-QwenPI model. This work demonstrates that sufficiently high-fidelity human data can remove the need for costly real-robot teleoperation in the final stages of policy training.
Source: HuggingFace Papers

Audit of GPQA and MMLU-Pro Reveals 12% Broken Questions, New Clean Versions Released
An audit of GPQA, MMLU-Pro, and MMMU-Pro benchmarks revealed that up to 12% of questions were broken and required removal.
An audit of major AI benchmarks including GPQA, MMLU-Pro, and MMMU-Pro found that up to 12% of the evaluation questions were broken and required removal. These benchmarks are standard tools for measuring the reasoning and knowledge capabilities of large language models, making the integrity of their data sets essential for accurate performance tracking. The identification of these flawed questions suggests that current LLM performance rankings may have been slightly skewed by inaccuracies in the underlying test sets. Researchers have consequently released updated "clean" versions of these datasets to ensure more rigorous and reliable model comparisons moving forward. This development highlights a critical trend in the AI research community to scrutinize the quality of evaluation data as much as the models themselves. By removing ambiguous or incorrect questions, developers can obtain a clearer picture of actual model performance improvements across various domains.
Source: r/LocalLLaMA
AI Applications
Applications that matter are domain-shaped: Andrew Ng’s LearnVector bets on guarded one-to-one tutoring, while medical robotics leans on GPU-native physics sim to fill data gaps you cannot scrape from the public internet.
Andrew Ng Launches LearnVector to Scale One-to-One AI Tutoring
$100M investment from Coursera
We're heads‑down building, and will have products to show by early 2027.
Andrew Ng has launched LearnVector with a $100 million investment from Coursera to transition educational models from one-to-many instruction to personalized one-to-one learning. The company addresses the limitations of current chatbots, which often facilitate cognitive offloading and harm actual skill acquisition by providing quick answers without pedagogical guardrails. By leveraging advanced AI, the platform aims to create a trustworthy learning guide that plans custom paths, adapts to individual styles, and ensures mastery through patient engagement. LearnVector intends to collaborate with established platforms like Coursera and Udemy to utilize their libraries of authoritative materials for training its systems. The venture focuses on making high-quality tutoring economically accessible to everyone, aiming to solve the historical scarcity of personalized teaching. Products from the new company are expected to be publicly showcased by early 2027 as the team focuses on foundational development.
Source: Hacker News

Advancing Healthcare Robotics via GPU-Native Medical Physics Simulation
healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation
Every demonstration requires specialized equipment, clinical expertise, and access to patients or laboratory environments.
Healthcare robotics development faces significant hurdles because it cannot leverage internet-scale data or unlimited real-world experimentation compared to autonomous driving. Developers encounter a primary data gap where every demonstration necessitates specialized equipment, clinical expertise, and access to actual patient environments. GPU-native medical physics simulation addresses these challenges by creating high-fidelity virtual environments for training robotic policies. This approach allows for the generation of synthetic data that mimics complex human anatomy and medical procedures without the risks associated with real-world clinical trials. By bridging the gap between simulation and reality, NVIDIA's tools enable more rapid iteration and testing of medical robotic systems. The technology effectively lowers the barrier to entry for developing sophisticated robotic interventions in clinical settings while maintaining safety and precision.
Source: NVIDIA Generative AI Blog
AI Policy & Ethics
Frontier employees are petitioning the US government to manage development pace—another signal that capability jumps (crypto, agents, overnight burn) are now treated as public risk, not just product velocity.
1,100 AI Employees Petition US Government to Regulate Frontier AI Development Pace
1,100 current/former frontier-AI employees sign a petition calling for US gov't to step in
A group of 1,100 current and former employees from frontier AI developers has signed a collective petition calling for the United States government to intervene in the development process. The petition specifically requests government action to manage the "pacing" of frontier AI development to ensure it proceeds at a safe and controllable rate. This initiative reflects a significant internal push from technical professionals for formal federal oversight rather than relying solely on industry self-regulation. By urging the government to step in, these signatories signal a concern that current development speeds may outpace existing safety frameworks. The collective signature count represents a substantial portion of the specialized workforce engaged in high-level AI research and deployment. This movement aims to shift the responsibility of risk management from private corporations to national regulatory bodies.
Source: r/LocalLLaMA
AI Infrastructure
Production agents often pass evals and fail in the wild. Real-time scoring and drift surfaces are becoming as important as quota dashboards.
Prefactor: Real-Time Evaluation Layer for AI Agents
We score every agent run in real time, surface quality regressions and drift as they happen
Most agents pass their evals and fail in production. Prefactor is the evaluation layer that closes the gap.
Prefactor scores every AI agent run in real time to bridge the performance gap between controlled evaluations and live production environments. The platform specifically targets engineering teams shipping agents to customers by surfacing quality regressions and performance drift as they occur at scale. By providing a dedicated evaluation layer, it allows developers to observe how their agents behave beyond static benchmarks and controlled testing scenarios. This observability tool helps teams identify exactly where agents are failing during live interactions, offering deep visibility into agent behavior. The system is designed to provide actionable insights into agent performance, ensuring that models maintain reliability once deployed in the wild. Users can monitor metrics across all agent runs to maintain high-quality outputs and minimize unexpected behavior in real-world applications.
Source: Product Hunt
This report is generated by WindFlash AI with manual editorial focus, based on public AI news from the past 48 hours.