The infrastructure for artificial intelligence is fracturing into layers of control, cost, and behavior, with the margin battle shifting from model capability to who manages the decision point. Amazon is dismantling nondisclosure agreements around data centers while SoftBank outsources its AI ambitions to DigitalBridge, a signal that data center economics have become the binding constraint on deployment. Simultaneously, the commodity layer is collapsing: Aleph Alpha shipped Kolibri as an open-weight 78.1 billion parameter mixture-of-experts model that activates only 3.46 billion parameters per token, DeepSeek released desktop agent applications under MIT licensing, IBM made Bob self-hostable for on-premises deployment, and Microsoft launched MAI-Transcribe-2-Streaming at $0.54 per hour. The actual competition has moved downstream into inference serving and agent behavior, where Prime Intellect runs on Blackwell and NVIDIA's DGX Spark targets $199 desktops. What emerges is not competition between models but competition over who controls the decision layer and the right to interrupt your attention.
Yet the infrastructure layer and the commodity layer are exposing a production reliability crisis that benchmarks and leaderboards are not measuring. Hugging Face named a fundamental problem: agents declare task completion while underlying databases show otherwise, a gap where real money gets lost in production environments. The SWE-rebench leaderboard shows complete stability at the top tier, with AnthropicFable 5 holding 64.5 percent, identical to the previous cycle, suggesting either saturation or stalled incremental progress. Database research reveals that relational systems are being retrofitted to serve AI workloads while confronting limits in classical design, with semantic data integration and agentic workflows introducing new operational requirements like branching, versioning, and governed memory. Testing infrastructure like DBcover and ERIQ are detecting logic bugs that binary accuracy metrics miss entirely, yet these tools remain niche while the industry markets agentic AI as solved.
GitHub repositories tell the story of where actual production pressure lives. Token optimization dominates the conversation: Caveman cuts token usage by 65 percent, context-mode achieves 98 percent output sandboxing efficiency. Memory and context persistence have become table stakes through compression and selective injection, reflecting a shift from proof-of-concept to systems that must run repeatedly without retraining on full context. Infrastructure plays like Cloudflare's agent workspace and Kiln's evaluation platform indicate developers are asking how to operationalize agents within existing systems, not whether agents can code. The repositories that gain traction solve one problem well and integrate with existing tools rather than attempting to be comprehensive frameworks. The social cost surfaces in parallel: OpenAI lost a safety employee over broken culture, Meta's Muse collects detailed profiles at millions of downloads, and the price of anything with memory spikes because the entire commodity chain has been conscripted into AI infrastructure. What emerges is not a solved problem but a field racing to deploy systems it does not yet know how to keep reliable.
Grant Calloway
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
Relational foundation models (RFMs) are pretrained once on a collection of relational databases and prediction tasks, and then applied zero-shot to previously unseen databases and tasks. To make a prediction for a target row, an RFM samples a neighborhood of rows linked to that row through foreign keys and uses this neighborhood as its inference context. Lowering inference cost is an important goal for any foundation model, and for RFMs this cost grows with the size of the context. The simplest ways to shrink the context is to drop some of the sampled rows, but this ignores the semantics of the database schema, so it is as likely to discard informative rows as uninformative ones. We propose STEER, a sampling approach that shrinks the inference context by concentrating it on the tables most relevant to the prediction task at hand. STEER obtains relevance information by prompting a large language model to rank the foreign-key edges of the database schema into relevance tiers for the given task, and then maps each tier to a probability of following that edge during traversal. Because the ranking uses only the schema, it is computed once per task and reused across all subsequent predictions, amortizing its cost. We evaluate STEER on three state-of-the-art RFMs (RT, RT-J, and Griffin) and show that it reduces inference context size by about 40% on average while maintaining, and in some cases improving, prediction accuracy.
Reconfigurable Intelligent Surfaces (RIS) are emerging as a key technology for programmable wireless environments in the beyond the fifth generation (B5G) networks. However, data-driven RIS research remains bottleneck by the lack of standardized, high-fidelity and open-source datasets. In this paper, we introduce a large-scale 3GPP TR 38.901-compliant dataset for RIS-aided millimeter wave (mmWave) networks, that considers severe path loss, blockage sensitivity, and spatial channel sparsity make the RIS assistance more impactful. The dataset spans various canonical 3GPP deployment scenarios across 20 controlled variants, capturing diverse user densities, fading conditions, and blockage regimes. Uniquely, every sample includes oracle RIS phase configurations obtained via a globally optimal brute-force codebook search, providing gold-standard supervision labels that are absent from any existing public dataset. Rich multi-task annotations comprising full channel state information (CSI), per-link channel decomposition, optimal phase matrices, and channel quality index (CQI) labels support a broad range of machine learning paradigms and downstream tasks, including phase optimization, channel estimation, and interference management. As the primary benchmark task, we introduce a novel CSI-to-CQI mapping that frames RIS-aided link-quality prediction as a scalable scalar classification problem, thereby avoiding the exponential output complexity of the direct phase vector prediction. We have evaluated this mapping against state-of-the-art architectures under in-distribution, out-of-distribution, and real-world hardware measurement conditions. Our dataset provides a reproducible, extensible, and community-ready foundation to accelerate data-driven research in RIS-aided B5G networks.
Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. However, LLM-generated plans are error-prone: a plan may silently compute something other than what was asked, fail during execution, or return a result that misses the question. This paper presents WeaveData, a multimodal data analysis system with self-critiquing and self-evolving LLM plans. First, WeaveData generates a typed logical plan for each question and critiques it step by step before execution, and it checks the executed result against the question afterwards. Second, WeaveData evolves a plan that fails or misses the question: it diagnoses the failure with the actual data, reuses the results that remain valid, and accumulates planning experience for later questions. Third, WeaveData grounds planning in a metadata knowledge graph of all modalities, clarifies ambiguous questions with the user, and backs every model judgment with evidence in an interactive notebook. We demonstrate WeaveData on two public multimodal datasets.
We design, implement, and evaluate KathDB-FAO, a new query evaluation subsystem for our KathDB multimodal DBMS. KathDB-FAO takes as input a query in natural language (NL) and converts it into a query execution plan where each operator is a function whose body is synthesized during query evaluation, which allows powerful query-specific optimizations. To generate accurate and efficient plans from NL, KathDB-FAO first extracts fine-grained atomic actions for correctness, then establishes contracts on the inputs and outputs of those actions and groups them for efficiency, and finally synthesizes the function for each group on the fly. On SemBench, KathDB-FAO cuts execution cost by 58.8% on average across scenarios compared with the next best system, at comparable or better quality.
Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source--memory integration, through which enterprise sources and accumulated memory can be utilized together, and memory governance, through which shared memory remains subject to organizational policies throughout its lifecycle. These requirements interact when information from enterprise sources persists in memory. As this information is repeatedly derived and reused under changing principals and policies, source restrictions may be bypassed, resulting in information leakage. Preventing such leakage requires authorization continuity, under which source restrictions remain effective throughout source-to-memory and memory-to-memory derivation and reuse. Existing approaches address these concerns individually, but do not treat source--memory integration, memory governance, and authorization continuity as combined core design targets across the memory lifecycle. We define Governed Enterprise Memory as agent memory designed around this combined scope and present AkasicMEM as its realization. AkasicMEM realizes authorization continuity through transitive lineage, policy composition during memory formation, and policy re-evaluation during retrieval. It is built on GraphAI's AkasicDB, a unified vector--graph--relational database whose storage and execution substrate enables the underlying operations of these mechanisms to be jointly optimized and executed.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 57.6 | 97 | $8.00 |
| 2 | Claude Sonnet 5.5 | 56 | 138 | $4.00 |
| 3 | Claude Fable 5.1 | 53.4 | 69 | $20.00 |
| 4 | GPT-6 Astra | 52.7 | 63 | $20.00 |
| 5 | Gemini 4 Argon | 52.6 | 0 | $4.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | AnthropicFable 5 [high]Model | 64.5%± 1.41% |
| 2 | GrokGrok 4.5 [high]Model | 63.8%± 0.60% |
| 3 | AnthropicOpus 5 [high]Model | 63.4%± 1.35% |
| 4 | Z.aiGLM-5.2 [high]Model | 62.9%± 1.19% |
| 5 | OpenAIGPT-5.6 Sol [medium]Model | 62.3%± 1.83% |
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
The design language that makes your AI harness better at design.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Build production-ready applications in TypeScript
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
🐍 Geometric Computer Vision Library for Spatial AI
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Infrastructure for continually self‑improving agents
Close the lid and walk away — every window, desktop and terminal on your machines, live in any browser. E2E-encrypted, no open ports, no client, no account.
Paxeer X is a Distributed HyperState Machine for payments, code execution and intent routing. Designed for Machines and the Users operating them.