Capital and technical capability have outpaced institutional agreement on control. Anthropic maintains boundaries around its models despite Pentagon pressure because it has consumer revenue to rely on. Hugging Face released OpenClaw to position open tooling as a counterweight to proprietary lock-in. Meanwhile, the infrastructure spreads faster than governance can follow: AI-driven border surveillance rolls out across West Africa with minimal oversight, recruiters work around AI hiring systems because they've lost confidence in outputs, and xAI has hemorrhaged nearly all its co-founders, suggesting no consensus exists even within single companies on what the technology should optimize for.
Research in computer science and society reveals why this fragmentation persists. Studies on AI in education show epistemic authority and accountability cannot be treated as technical problems alone. Work on fairness exposes how aggregate metrics systematically fail to capture real harms, producing false assurance about system safety. Structural asymmetries of power and accountability inhere in system design itself and cannot be remedied by transparency or consent alone. The research reframes AI governance not as a problem of better alignment or clearer rules, but as a problem of institutional and technical architecture: who decides, who bears consequences, and whether those roles can be reconciled within a single interaction.
On the development side, the pattern is clear: teams have stopped asking whether agents work and started asking how to make them reliable and auditable. Trending repositories show the industry moving past hype into operational concerns. AgentScope's explicit pitch around visibility and trust reflects this shift. Developers are solving last-mile problems, adapting general models to specific hardware and tasks without cloud dependencies, and building specification systems that can be versioned and reasoned about rather than relying on prompt engineering. The unglamorous work of making agents useful enough to deploy is underway. What remains unresolved is whether any company, government, or standards body can enforce a coherent answer to what these systems should optimize for when they operate across borders and jurisdictions with zero coordination.
Grant Calloway
The ongoing digital transformation of work, administration, health, mobility, and social interaction is profoundly reshaping everyday life, steadily shifting control from individuals to large platform providers. Although data is often labeled the "gold of the 21st century", its real value is realized through services that access, combine, and exploit it. Today, individuals have little sovereignty: life events (e.g., changing an address, insurance, job, or marital status) require fragmented, repetitive interactions across numerous systems, leaving users overwhelmed rather than empowered. We argue for a fundamental rethinking of this situation and propose a virtual, trustworthy platform, called "MyVirtualME" that acts on behalf of the human individual towards companies, administrations, and other actors. This platform centers services, data, permissions, and data usage around humans, not foreign actors, allowing genuine control and transparency. Our vision goes beyond a pure data focus: supported by trusted services ranging from basic notifications to intelligent assistants, MyVirtualME aims to reduce digital bureaucracy, improve and even automate interactions, and support informed oversight of an individual's health, financial, and administrative situation. We firmly believe that a human-centered MyVirtualME is essential for restoring personal sovereignty in our digital future.
AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer "digital twins." Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).
This study introduces a dual-matrix computational architecture to mathematically quantify the morphological and theological divergence of 196 Hindu and Vajrayana Buddhist esoteric deities. Physical morphology is evaluated via a discrete Gower distance matrix enhanced by a novel "Cardinality Weighting" algorithm, while theological function is mapped via dense vector embeddings generated from Large Language Model (LLM) semantic expansions, explicitly utilized as a synthetic proxy to mitigate circular reasoning. The multi-modal topological projections provide algorithmic validation of "iconographic camouflage", demonstrating how distinct visual forms structurally obscure shared cross-tradition functions. Furthermore, I computationally model the "Atin Effect" - serving simultaneously as a psychological observation of sequential cognitive bias and a machine learning benchmark - demonstrating how high-cardinality esoteric anchors (e.g., a veena or a severed head) override systemic theological disparities to mathematically cluster orthodox and Tantric entities. Cross-tradition spatial analysis establishes that the highest esoteric manifestations, such as the Hindu Chinnamasta and the Buddhist Chinnamunda, share a near-identical mathematical coordinate across both visual ($D_G = 0.288$) and semantic ($D_C = 0.068$) boundaries, indicating a 1:1 esoteric transfer. By open-sourcing this architecture, I provide a scalable, unsupervised machine learning tool for Digital Humanities scholars and comparative theologians to rigorously map latent structural continuities across qualitative cultural corpora.
Educational technology (EdTech) platforms collect highly sensitive student data, including behavioral logs, disability records, and academic histories. However, privacy considerations are often postponed rather than treated as a foundational design requirement. We present a mixed-methods study combining 12 semi-structured interviews with EdTech professionals and a privacy policy audit of 48 platforms coded across five dimensions, with strong inter-rater reliability (mean Cohen's Kappa = 0.781). Our interviews reveal a recurring organizational pattern in which privacy is recognized as important but deferred across the product lifecycle as organizations prioritize product functionality, growth, funding, and immediate educational outcomes. Responsibility is often delegated to cloud providers, policy documents, or downstream institutions, while limited privacy-related feedback gives organizations little pressure to change these practices. The policy analysis reflects these patterns: platforms describe what data they collect relatively well but provide substantially less information about how that data is subsequently governed. Thirty-three percent make no meaningful Artificial Intelligence (AI) disclosure despite visible AI features, and 73% provide only generic accountability and breach-response language. K-12 platforms perform better on children's consent where regulation creates explicit requirements, but this advantage does not extend to AI governance or accountability. These findings suggest that meaningful improvement requires enforceable institutional and regulatory mechanisms rather than voluntary privacy commitments alone.
The rising global prevalence of mental health conditions, together with longstanding barriers in traditional healthcare, such as limited resources, high cost, stigma, and privacy concerns, has created an urgent need for accessible and scalable support. Large Language Models (LLMs) have emerged as a transformative technology with strong potential to democratize mental health support through advanced natural language understanding and generation. However, the rapidly expanding, fragmented body of work in this area lacks a coherent evolutionary narrative, making it difficult to contextualize current progress and identify future directions. This survey addresses this gap by organizing and analyzing the literature around a central thesis: the role of LLMs in mental health is evolving through three distinct, increasingly sophisticated phases. We trace this trajectory from Phase I, in which LLMs act primarily as passive Information Tools and Pattern Recognizers for assessment; through Phase II, where they function as Empathetic Conversationalists for in-the-moment, stateless interactions; to the current frontier, Phase III, which seeks Longitudinal, Personalized Companions implemented as stateful cognitive agents. To support this framework, we systematically review core technologies, agent architectures (Profile, Memory, Reasoning, and Planning), and the critical infrastructure of datasets and benchmarks, highlighting how their evolution underpins this developmental path. Viewing the field through this developmental lens, we provide a comprehensive synthesis of existing work, an insightful narrative of its trajectory, and a clear roadmap for future innovation in responsible, effective, and human-centered AI for mental healthcare. A curated collection of the resources reviewed in this survey is available at our project repository: https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.
The need for collaboration between diverse fields of research is increasingly recognised as important by research funding agencies. A significant driver of this need is the current revolution in artificial intelligence (AI) and related technologies. There is a growing interest in the potential impact of AI in different fields including the methodologies they use and the resulting advances in new knowledge, new access and enhanced productivity. However, there is also a corresponding increase in concern about the fundamentals of AI technologies and the way in which trans and/or interdisciplinary research is approached. The resulting collaboration too often ends up as a one-way street where the domain partner acts only as an information provider. For example, the contribution of the AHSS partner might be limited to providing insight about ethics and/or the technology partner may only provide a service to build applied AI-based solutions. In response to this problem, we propose a reciprocal approach to collaboration where both partners seek to understand, cooperate and identify jointly significant impacts. In this paper we explore this relationship between cultural heritage institutions (GLAMs), Arts, Humanities & Social Sciences (AHSS) research and technology-led AI research, especially the impact of current technological advances in AI. Drawing from the history of convergence in GLAM studies, we propose five key practices to form a framework for greater understanding across this divide.
Composite score across coding, math, and reasoning
| # | Model | Score | tok/s | $/1M |
|---|---|---|---|---|
| 1 | GPT-5.4 | 57.2 | 84 | $5.63 |
| 2 | Gemini 3.1 Pro Preview | 57.2 | 115 | $4.50 |
| 3 | GPT-5.3 Codex | 54 | 81 | $4.81 |
| 4 | Claude Opus 4.6 | 53 | 56 | $10.00 |
| 5 | Claude Sonnet 4.6 | 51.7 | 66 | $6.00 |
Agentic coding on real-world software engineering tasks
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 65.3% |
| 2 | gpt-5.2-2025-12-11-medium | 64.4% |
| 3 | GLM-5 | 62.8% |
| 4 | gpt-5.4-2026-03-05-medium | 62.8% |
| 5 | Gemini 3.1 Pro Preview | 62.3% |
real time face swap and one-click video deepfake with only a single image
An agentic skills framework & software development methodology that works.
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
An autonomous agent for deep financial research
Building a modern alternative to Salesforce, powered by the community.
Minimalist web-searching platform with an AI assistant that runs directly from your browser. Uses WebLLM, Wllama and SearXNG. Demo: https://felladrin-minisearch.hf.space
∞ Generate endless answers from all-knowing ChatGPT (on any topic!)
MySQL-compatible HTAP database with Git for Data, vector search, and fulltext search. Cloud-native, AI-ready
Train Large Language Models on MLX.
Tensors and Dynamic neural networks in Python with strong GPU acceleration