Meanwhile, the AI industry is fracturing along a new fault line between those building for government and those refusing to. Anthropic's lawsuit against the Pentagon over its supply-chain-risk designation has drawn public support from more than 30 OpenAI and Google DeepMind employees signing an amicus brief, yet OpenAI itself has moved in the opposite direction, acquiring Promptfoo to strengthen its ability to deploy AI agents in critical operations while Caitlin Kalinowski, the company's head of robotics, resigned over inadequate safeguards in its Pentagon contract. The split reflects a deeper disagreement about what builders should accept in exchange for scale and legitimacy, with the designation already costing Anthropic material revenue as companies paused deal talks.
The money is flowing toward infrastructure and specialized models rather than consolidation around any single foundation. Yann LeCun's AMI Labs closed a $1.03 billion seed round at a $3.5 billion valuation to build world models focused on physical understanding, while Nscale, an Nvidia-backed infrastructure startup, reached a $14.6 billion valuation on a $2 billion raise. Anthropic launched a Claude Marketplace to streamline enterprise procurement and deployed Code Review, a multi-agent system for analyzing AI-generated code. The market is settling into layers: frontier model providers compete on capability and trust, infrastructure companies capture deployment economics, and specialized tools fill gaps between raw models and production use.
The practical pressure on builders is now acute. Amazon held an engineering meeting after AI-related outages linked to generative AI-assisted code changes, while Microsoft's Copilot for Microsoft 365 has captured only 3 percent of its customer base despite two years in market, forcing the company to add Anthropic's Claude to its own tools. The market is no longer asking whether AI works in theory but whether it works reliably enough to deploy at scale, whether it can be audited and reviewed, and whether users can understand what it does. Lab announcements reveal a hardening focus on the operational layer: security, observability, and cost reduction in production environments. The companies that win will solve these problems faster than their competitors, not those that promise the most capability.
Grant Calloway
Social media recommendation feeds often optimize for users' immediate impulses rather than preferences they would hold after deeper reflection. Some systems address this misalignment by incorporating users' explicit preferences via a configuration page or in-feed controls instead of just behavioral signals. However, users typically have evolving preferences, and their stated preferences and behavior naturally diverge, necessitating continuous reflection and feed realignment. But existing strategies require the user to take initiative and are often effortful; as a result, in practice they are rarely invoked. We present Compass, a system that aligns a user's feed with their reflective preferences by helping users reflect on and articulate their preferences given their behavior. To enable continuous reflection during everyday browsing, Compass surfaces in-situ reflections via lightweight notifications, while feed alignment is achieved by periodically simulating behavioral signals and directly manipulating feed content. We embedded Compass within YouTube Shorts and compared it against a baseline without continuous support through a 10-day field study (N=15). We found that Compass promoted more reflective and purposeful feed consumption, iterative preference adjustment, and stronger feed alignment, without sacrificing the casual nature of feed browsing.
High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images. Current AI-assisted annotation tools often lack assistance or rely on one-way workflows where experts have to perform extra manual calibrations to improve AI models, resulting in limited efficiency. To address this, we propose Bidirectional Human-AI Augmentation(BiHAA), a closed-loop framework in which skills and domain knowledge base evolve through real-time interaction and bidirectional HAI augmentation. Informed by a formative study with 20 artwork annotators from different backgrounds, we implement this framework in ArtAnno, an artwork annotation system driven by a multi-agent architecture. The system includes a Proactive Agentic Support Module, where AI augments humans through semantic mining and label suggestion, and an Interaction-Driven Evolution Module, where human expertise continuously enhances the AI through distilling annotation trajectories into reusable experience. Evaluation through a user study and two case studies demonstrates that our framework and system improve annotation efficiency, enable knowledge accumulation, and reduce the effort of information seeking and verification for annotators with limited domain expertise. We conclude by discussing broader implications and future directions.
This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits on breadboards and printed circuit boards (PCB) by undergraduates through conversations with public-domain large language models (LLMs). Through thematic analysis of students' voluntarily shared chat logs when debugging pre-determined buggy circuits under exam and time pressure, we discovered multimodal usage patterns by students and considerable domain knowledge and sensible debugging suggestions offered by off-the-shelf LLMs. Meanwhile, we also identified major gaps in LLM technologies and students' skills during human-AI collaborative debugging, such as LLMs' limitations in 2D/3D image-based reasoning, unjustified tone of confidence, and students' deficits in fundamental concepts and critical thinking.
Commercial wearable devices continuously capture rich physiological data (e.g., heart rate, respiration), opening new possibilities for monitoring health conditions, notably around stress. Despite their promise, turning raw wearable physiological data streams into visualizations that surface stress-related insights in daily activities, and that ultimately foster reflection, awareness, and better stress management, remains a significant challenge. The data are noisy and context-dependent: the same spike in heart rate can come from sprinting, a tense presentation, or laughing with friends. To address these challenges, we propose a framework that combines user annotations with wearable data to support better stress management. We introduce a web framework offering interactive visualizations that layer daily activities, stress events, and interventions onto raw physiological streams, enabling users to reflect and identify trends. In a four-week pilot with seven university graduate and undergraduate student participants who logged 269 events, our tool revealed patterns between different types of interventions and stress: social interaction reduced average heart rate by 4.35 to 5.0 beats per minute, deliberate rest reduced average Garmin stress scores by 10.03 to 13.83 points, and mindfulness activities decreased average HRV by 6.61 to 13.22 milliseconds.
Non-invasive blood glucose level (BGL) estimation from photoplethysmography (PPG) holds great promise for wearable health monitoring, but results across studies are hard to compare due to inconsistent datasets, data leakage, and non-standardized evaluation metrics. We present the first reproducible, extensible evaluation pipeline and use it to reassess five representative PPG-based BGL methods on published datasets under three increasingly strict data-split protocols: random window-level, participant-aware, and leave-some-participants-out (LSPO). Models appeared competitive under random splitting but collapsed under participant-aware and LSPO evaluation, with nearly all yielding near-zero or negative R$^2$ values comparable to a mean-prediction baseline. Critically, across every model and split, over 90% of predictions fell within clinically acceptable zones (Clarke Error Grid A+B), including the baseline. This reveals a fundamental disconnect: clinical zone metrics systematically conceal model failure in this domain. Our findings demonstrate that random train-test splits substantially overestimate the generalization of PPG-based BGL models due to sample-level data leakage, and that robust ML evaluation must precede clinical validation to meaningfully assess real-world utility.
As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become a fundamental behavioural question. We introduce a behavioural measurement framework combining intent and delegated decision authority to quantify what consumers seek from AI and how much decision-making authority they assign to it. Applied to 1.5 million real-world ChatGPT and Gemini interactions from 6,304 users in the United States and India, we find that financial services are already a substantial AI use case. Consumers overwhelmingly use AI to retrieve information and shape financial judgement, while delegation of financial execution remains rare. By shifting attention from conversation topics to delegated decision authority, this work establishes a behavioural baseline for measuring the transition to increasingly agentic AI.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Gemini 3.1 Pro Preview | 57.2 | 110 | $4.50 |
| 2 | GPT-5.4 | 57 | 78 | $5.63 |
| 3 | GPT-5.3 Codex | 54 | 68 | $4.81 |
| 4 | Claude Opus 4.6 | 53 | 55 | $10.00 |
| 5 | Claude Sonnet 4.6 | 51.7 | 69 | $6.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | Claude Code | 52.9% |
| 2 | Junie | 52.1% |
| 3 | Claude Opus 4.6 | 51.7% |
| 4 | gpt-5.2-2025-12-11-xhigh | 51.7% |
| 5 | gpt-5.2-2025-12-11-medium | 51.0% |
Sample code and notebooks for Generative AI on Google Cloud, with Gemini on Vertex AI
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
The best ChatGPT that $100 can buy.
The design language that makes your AI harness better at design.
Build resilient language agents as graphs.
OpenVision (ICCV 2025), OpenVision 2 (CVPR 2026), and OpenVision 3
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search, and conversational systems.
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
A lightweight, developer-focused database management tool. Supports MySQL, PostgreSQL and SQLite. Hackable with plugins. Built for speed, security, and aesthetics.