Today's news confirms what the market has already priced in: the AI industry's center of gravity has shifted from capability announcements to operational durability. Token prices are cooling and spending growth is plateauing, yet the actual competition is intensifying precisely where it matters least to the headlines. Meta is squeezing efficiency from recycled RAM through custom chips while managing unionization pressure at DeepMind. Cloud providers are fracturing under their own outages. Security patches are accelerating because vulnerabilities now arrive faster than fixes can ship. The companies winning aren't those with the biggest models; they're the ones solving the unglamorous problems of hardware reuse, infrastructure resilience, and exploit remediation.
AMD's three infrastructure announcements in a single day exemplify this shift. The company is not launching models or frameworks. It is publishing benchmarks that directly compare competing AI agents on the same hardware, building integrated pipelines that keep data in VRAM to eliminate bottlenecks, and optimizing performance for models already shipping from competitors. AMD names Cursor Agent, Claude Code, and OpenAI Codex in its coding benchmark. It names Kimi-K2.5 and MiniMax-M2.5 in inference optimization. The message to the market is explicit: your models will run faster on our hardware if you follow our patterns. This is the work of infrastructure positioning, not foundation model competition. Database research reinforces the pattern, clustering around three interconnected frontiers: agentic systems for data manipulation, semantic integration of unstructured and structured data, and the infrastructure required to make agents reliable at scale. Rather than treating agents as stateless compute or semantic operations as isolated problems, the emerging research treats the artifacts of search and the semantics of data as first-class database objects, queryable and governed by the infrastructure itself.
Benchmark stability and GitHub trends confirm where actual differentiation is occurring. OpenAI's gpt-5.5-2026-04-23-xhighModel holds the top position on SWE-rebench at 62.7 percent with tight confidence intervals, while Claude Fable 5 leads Artificial Analysis at 59.9 percent, suggesting different problem structures rather than a unified capability hierarchy. The trending repositories reveal a decisive shift toward agentic tools embedded into developer workflows: Claude Code dominates with the highest star count as a terminal-resident coding agent, while supporting infrastructure clusters around agent isolation, auditable memory systems, and skill standardization. The newcomers are not flashy consumer tools but unglamorous plumbing. Elasticsearch, PyTorch, and Supabase remain foundational, but TencentCloud's CubeSandbox isolates agent execution, shodh-memory offers local auditable agent memory without API dependencies, and academic-research-skills packages domain expertise as Claude Code skills. The pattern is consolidation around a specific vision of agency: agents as terminal-native, skill-based, auditable systems that integrate into existing infrastructure rather than replace it.
Grant Calloway
Reconfigurable Intelligent Surfaces (RIS) are emerging as a key technology for programmable wireless environments in the beyond the fifth generation (B5G) networks. However, data-driven RIS research remains bottleneck by the lack of standardized, high-fidelity and open-source datasets. In this paper, we introduce a large-scale 3GPP TR 38.901-compliant dataset for RIS-aided millimeter wave (mmWave) networks, that considers severe path loss, blockage sensitivity, and spatial channel sparsity make the RIS assistance more impactful. The dataset spans various canonical 3GPP deployment scenarios across 20 controlled variants, capturing diverse user densities, fading conditions, and blockage regimes. Uniquely, every sample includes oracle RIS phase configurations obtained via a globally optimal brute-force codebook search, providing gold-standard supervision labels that are absent from any existing public dataset. Rich multi-task annotations comprising full channel state information (CSI), per-link channel decomposition, optimal phase matrices, and channel quality index (CQI) labels support a broad range of machine learning paradigms and downstream tasks, including phase optimization, channel estimation, and interference management. As the primary benchmark task, we introduce a novel CSI-to-CQI mapping that frames RIS-aided link-quality prediction as a scalable scalar classification problem, thereby avoiding the exponential output complexity of the direct phase vector prediction. We have evaluated this mapping against state-of-the-art architectures under in-distribution, out-of-distribution, and real-world hardware measurement conditions. Our dataset provides a reproducible, extensible, and community-ready foundation to accelerate data-driven research in RIS-aided B5G networks.
Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. However, LLM-generated plans are error-prone: a plan may silently compute something other than what was asked, fail during execution, or return a result that misses the question. This paper presents WeaveData, a multimodal data analysis system with self-critiquing and self-evolving LLM plans. First, WeaveData generates a typed logical plan for each question and critiques it step by step before execution, and it checks the executed result against the question afterwards. Second, WeaveData evolves a plan that fails or misses the question: it diagnoses the failure with the actual data, reuses the results that remain valid, and accumulates planning experience for later questions. Third, WeaveData grounds planning in a metadata knowledge graph of all modalities, clarifies ambiguous questions with the user, and backs every model judgment with evidence in an interactive notebook. We demonstrate WeaveData on two public multimodal datasets.
We design, implement, and evaluate KathDB-FAO, a new query evaluation subsystem for our KathDB multimodal DBMS. KathDB-FAO takes as input a query in natural language (NL) and converts it into a query execution plan where each operator is a function whose body is synthesized during query evaluation, which allows powerful query-specific optimizations. To generate accurate and efficient plans from NL, KathDB-FAO first extracts fine-grained atomic actions for correctness, then establishes contracts on the inputs and outputs of those actions and groups them for efficiency, and finally synthesizes the function for each group on the fly. On SemBench, KathDB-FAO cuts execution cost by 58.8% on average across scenarios compared with the next best system, at comparable or better quality.
Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source--memory integration, through which enterprise sources and accumulated memory can be utilized together, and memory governance, through which shared memory remains subject to organizational policies throughout its lifecycle. These requirements interact when information from enterprise sources persists in memory. As this information is repeatedly derived and reused under changing principals and policies, source restrictions may be bypassed, resulting in information leakage. Preventing such leakage requires authorization continuity, under which source restrictions remain effective throughout source-to-memory and memory-to-memory derivation and reuse. Existing approaches address these concerns individually, but do not treat source--memory integration, memory governance, and authorization continuity as combined core design targets across the memory lifecycle. We define Governed Enterprise Memory as agent memory designed around this combined scope and present AkasicMEM as its realization. AkasicMEM realizes authorization continuity through transitive lineage, policy composition during memory formation, and policy re-evaluation during retrieval. It is built on GraphAI's AkasicDB, a unified vector--graph--relational database whose storage and execution substrate enables the underlying operations of these mechanisms to be jointly optimized and executed.
Graph databases are frequently positioned as categorically necessary for connected-data workloads, yet the systems dimension along which they actually differ - query planning, indexing, and data-readiness cost - is rarely isolated from vendor framing. We construct a synthetic, biomedical-shaped property graph (1.02 million nodes, 5.34 million total node and edge rows) and a twenty-query workload spanning neighborhood lookups, bounded paths, set intersections, anti-joins, grouped aggregation, top-k ranking, temporal filters, full scans, and relational joins. We benchmark Corvic AI - a purpose-built columnar query engine underlying Corvic's ontology management layer ("memories")- against seven purpose-built or graph-extension database systems (LoraDB, Ladybug, DuckPGQ, Memgraph, Neo4j, HugeGraph, and FalkorDB) at three graph scales spanning three orders of magnitude. We report query latency geomeans, bulk-ingest throughput, point-update latency, and answer correctness for each system, and we derive a simple total-cost-of-ownership model that expresses the ingest/query trade-off as a function of query volume. Our central finding is that no system in this sample is categorically fastest: a native graph engine (Ladybug) outperforms Corvic AI on narrow, bounded-neighborhood shapes, while Corvic AI is faster on shapes that scan or join a large fraction of the graph, and a system implementing graph query syntax via SQL/PGQ (DuckPGQ) is measurably slower purely due to query-plan choice. The dominant cost differential in our data is not query latency but the cost of making data queryable at all: bulk-ingest throughput varies by three orders of magnitude across engines (5.0k-4.3M rows/s), a gap that a simple crossover-point calculation shows dominates total cost for any workload with fewer than roughly 105 queries per data refresh.
Traditional data systems face profound limitations in the AI era, relying on human-crafted pipelines, lacking semantic understanding of heterogeneous data, and operating through rigid, reactive processing. To address these challenges, we propose a new paradigm called the Data Agent, designed to manage, process, and analyze data with minimal human intervention. Data agents autonomously execute a wide range of data-related tasks, transforming traditional data systems by shifting from manual design to autonomous orchestration, from literal manipulation to semantic interpretation, and from reactive to proactive processing. Our Data Agent system includes six components: semantic data organization, semantic operators, agentic pipeline orchestration and optimization, feedback-driven refinement, memory management, and proactive adaptation. Building on this foundation, we also develop two specialized agents: the data analytics agent and the data science agent. Experiments on real benchmarks demonstrate significant performance gains of our data agent over state-of-the-art methods. We identify open challenges to guide future research in building fully autonomous data systems.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 59.9 | 62 | $20.00 |
| 2 | Claude Opus 4.8 | 55.7 | 61 | $10.00 |
| 3 | GPT-5.5 | 54.8 | 88 | $11.25 |
| 4 | Claude Opus 4.7 | 53.5 | 49 | $10.00 |
| 5 | Claude Sonnet 5 | 53.4 | 79 | $6.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | OpenAIgpt-5.5-2026-04-23-xhighModel | 62.7%± 0.91% |
| 2 | JunieJunieAgent | 61.6%± 0.64% |
| 3 | OpenAICodexAgent | 60.4%± 1.37% |
| 4 | AnthropicClaude CodeAgent | 59.6%± 1.98% |
| 5 | OpenAIgpt-5.5-2026-04-23-mediumModel | 58.9%± 0.78% |
Open-source AI hackers to find and fix your app’s vulnerabilities.
Use Codex from Claude Code to review code or delegate tasks.
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
Free and Open Source, Distributed, RESTful Search Engine
Action for checking out a repo
VisualTorch aims to help visualize Torch-based neural network architectures.
autoupdate paper list
High-Performance Symbolic Regression in Python and Julia
Cognitive brain for Claude, AI agents & edge devices — learns with use, runs offline, single binary. Neuroscience-grounded 3-tier architecture with Hebbian learning.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale