Swissi AI Journal
The Swiss flag flying in front of a snow-covered mountain.

ISSN 3043-1921

Swissi AI Journal Open Access & Peer-Reviewed

Artificial intelligence research from method to deployed system. Editors assess the method, evaluation and evidence, then publish accepted work continuously with open access from its publication date.

Open access

Peer-Reviewed Swissi Articles

Show all Swissi articles

SAIJ-5kdnql4rsq27Jul 2026

Multi-Jurisdictional Legal Identity Assurance for Capability Gating

A Design-Science Proposal for Tiered, Reusable Identity Assurance of Natural, Juridical, and Machine Entities

DOI 10.5281/zenodo.21901241

Flat maximum verification charges every participant for the rarest high-risk case, and excludes those who cannot clear a bar they never needed to. This model holds the assurance state apart from the capability gate that consumes it, so identity demand follows the act and the weight of its consequences rather than mere presence.

Keywords:
  • identity assurance
  • capability gating
  • tiered and reusable verification
  • multi-jurisdictional identity
  • data minimisation
  • entity taxonomy
  • bitemporal reliance
  • design science research

SAIJ-zpd6gvtrfaunJul 2026

Credentials and Triangulated Trust Signals on a Single Accountable Identifier

A Hash-Anchored Distributed-Ledger Framework for Portable Identity across Jurisdictions

DOI 10.5281/zenodo.21901243

Digital identity stays rigid while it is bound to provider accounts, mutable handles and local wallet schemes. An accountable hash-anchor tier sits above them: an inert root anchor, unlinkable profile anchors for distinct contexts, and gate-specific assurance evaluated at a point in time.

Keywords:
  • digital identity
  • verifiable credentials
  • accountable pseudonymity
  • distributed ledger
  • selective disclosure
  • identity assurance
  • self-sovereign identity
  • hash anchor

SAIJ-sh27g6sykt2kJul 2026

Identity-Staked Consensus and Collusion Resistance in Chartered Validator Sets

A Trust Model for Decentralised and Compliant Distributed Settlement Infrastructure

DOI 10.5281/zenodo.21901245

Permissioned ledgers are commonly dismissed as centralised because admission is restricted. Separating permissioning from control distribution makes validator identity externally costly collateral: public legal identity, charter state, liability and audit exposure, with affiliation-aware voting caps and per-member collusion margins.

Keywords:
  • identity-staked consensus
  • permissioned ledger
  • proof-of-authority
  • validator trust model
  • collusion resistance
  • settlement infrastructure
  • actor assurance
  • threshold class coverage
  • ledger evidence record

SAIJ-f3jtignfyge3May 2026

Firm Valuation When AI Shapes the Business Model

A Milestone-Based Real-Options Framework for the AI Valuation Uncertainty Problem

DOI 10.5281/zenodo.21901247

arXiv 2609.24181

Discounted cash flow, the IDW S 1 income approach and market multiples compress milestone probabilities, continuation options and risk shifts into opaque aggregate parameters. A milestone-gated real-options overlay decomposes that value into auditable components, with a Success Readiness Index deriving per-option probabilities from structured pairwise comparisons.

Keywords:
  • firm valuation
  • AI integration
  • real options
  • milestone-based valuation
  • intangible assets
  • AHP
  • multi-criteria decision analysis

SAIJ-cwo7xrcdsautMar 2026

Functional Architecture of European Electricity Trading Markets

Requirements for AI Supported Trading Systems under Regulatory Constraints

DOI 10.5281/zenodo.21901249

European electricity trading runs as a constrained multi-layer system in which legal design, exchange microstructure and network physics execute jointly across forward, day-ahead, intraday and balancing horizons. The paper specifies an AI-supported trading architecture with a permission gate on executable actions and fail-closed control logic under REMIT, MiFID II, MiFIR and EMIR.

Keywords:
  • EU electricity market
  • market coupling
  • NEMO topology
  • electricity balancing
  • AI trading systems
  • compliance-by-design

SAIJ-xz3bi3q7fwimAug 2025

Compliant AI Infrastructure for Regulated Finance

A tiered multi-agent framework with DLT audit trails for financial operations in DACH

DOI 10.5281/zenodo.21901251

arXiv 2609.27632

Regulation is treated as an orientation layer rather than a deterministic ruleset: a matrix of regulatory intent and exposure is compiled into concrete prohibitions, obligations and runtime budgets. Evidence, decisions and reason codes bind to a permissioned DAG, so a supervisor can replay how an outcome was reached and attribute failure.

Keywords:
  • DACH finance
  • regulated financial institutions
  • multi agent expert system
  • policy compiled orchestration
  • objective under constraints
  • permissioned DLT
  • DAG timestamping
  • audit trails
  • EU AI Act
  • MiFID II
  • DORA
  • GDPR
  • human oversight
  • execution gating
  • ESG budgets
  • verification and assurance

SAIJ-ddkjais6s332Aug 2025

A regulatory-compliant AI and verification system for higher education under ESG-aligned constraints

DOI 10.5281/zenodo.21901253

Two linked components for higher education: a role-specific multi-agent framework for institutional operations, and a decentralised verification layer for audit, credential authentication and tamper-evident records. GDPR, the EU AI Act, EQF, ECTS and ESG directives are encoded as structural constraints rather than checked after the fact.

Keywords:
  • Regulatory technology
  • artificial intelligence in education
  • multi-agent AI systems
  • decentralised verification
  • academic tokenisation
  • GDPR compliance
  • EU AI Act
  • digital credential infrastructure
  • ESG governance
  • UniAI
  • UniDVS

SAIJ-zkihhbpahsbrAug 2025

Verifiable Federated AI Infrastructure

Swiss compliant federated AI DLT network using Nash equilibrium and ESG metrics

DOI 10.5281/zenodo.21901255

Centralised AI infrastructure scales, and collides with latency, auditability and energy constraints. The design separates centralised training from decentralised inference and storage across five node classes, tying a size-neutral availability floor to tiered rewards for service level, ESG performance and anti-concentration.

Keywords:
  • decentralised data centre
  • AI
  • Federated AI infrastructure
  • ESG
  • ESG-aware compute
  • Nash equilibrium
  • digital sovereignty
  • Swiss data regulation
  • tokenised infrastructure
  • verifiable AI services

SAIJ-soeptiqyucowAug 2025

Generic AI-DLT Enterprise System

Architecture and methodology for scalable domain adaptation from a unified core framework

DOI 10.5281/zenodo.21901257

Compliance in AI deployments is usually applied afterwards, through prompt engineering, rather than built into the foundation. This architecture encodes regulatory, governance and ESG requirements as an objective-under-constraints problem, so every specialised agent operates within legally admissible and auditable bounds before any domain work begins.

Keywords:
  • Compliance-first AI
  • Multi-agent systems
  • Distributed ledger technology
  • Directed acyclic graph
  • Regulation by design
  • ESG integration
  • Domain-agnostic architecture
  • Deployment-agnostic architecture
  • Vendor-agnostic architecture
  • Objective-under-constraints

SAIJ-qzvrl4bwy7y2May 2025

Multi-Agent AI Architecture for Regulated Insurers

A generic AI framework under Solvency II and the AI Act in Austria and Germany

DOI 10.5281/zenodo.21901259

arXiv 2609.27636

The insurer is modelled as a constrained optimisation entity under solvency, legal, ESG and operational boundaries, then decomposed into specialised agents for capital, underwriting, claims, compliance and fraud. Human-in-the-loop roles enter through tiered access control, with an orchestrator enforcing regulatory admissibility across the set.

Keywords:
  • Multi-Agent Systems
  • Enterprise AI
  • Insurance Firms
  • Solvency II
  • AI Act
  • Regulated Environments
  • Constrained Optimisation
  • Principal-Agent Theory
  • Nash Equilibrium
  • Arrow’s Risk Pooling
  • Austria
  • Germany
  • Institutional Design
  • ESG Compliance
  • Regulatory Architecture
  • Model Context Protocol (MCP)
  • Agent-to-Agent Protocol (A2A)
  • AI Governance
  • Algorithmic Accountability
  • Financial Regulation

Standing calls

Current calls for papers

Standing calls identify current editorial priorities within the journal's full scope.

SustainabilityArtificial intelligence and ESG evidence
Results that cost less compute to obtain, artificial intelligence applied to generation and grids, and sustainability reporting treated as a question of evidence.
SovereigntyEnterprise AI under European compliance
Systems that run inside the jurisdiction: self-hosted or European inference, data governed within its required location, and measured capability and financial costs.
LanguageFoundation models for the DACH region
Models trained, adapted and evaluated on German-language material, with benchmarks that exist in German, and a clear account of where a regional model outperforms a general one.

From arXiv

Recent arXiv Articles

arXiv cs.AI

Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse

Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This...

arXiv cs.LG

The Drift Contract: Spectral Updates for Depth-Robust Local Learning

Local learning trains each layer with its own auxiliary loss and no global backward pass, which makes layer updates structurally parallel. Two problems have kept it marginal: accuracy degrades as depth grows, and hyperparameters are fragile. We apply Muon-style spectral update geometry (momentum orthogonalization...

arXiv cs.AI

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Modern language-model agents are built around the agent loop, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the...

arXiv cs.LG

Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection

Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit complex temporal structure, substantial acquisition variability, and a strong need for transparent decision-making. Although deep models can achieve high detection performance, they often provide limited insight into...

arXiv cs.AI

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers...

arXiv cs.LG

HARN: Hierarchical Associative Resonance Network for Event-Driven Multi-Timeframe Forecasting

Financial time series evolve across multiple temporal resolutions, challenging forecasting systems to incorporate newly available information without repeatedly recomputing unchanged representations. We introduce HARN, a Hierarchical Associative Resonance Network for event-driven multi-timeframe forecasting. HARN...

arXiv cs.AI

Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations

The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational...

arXiv cs.LG

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be...

arXiv cs.AI

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a...

arXiv cs.LG

LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels

Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deployment settings, however, labels come from weak annotators, historical decisions, heuristics, or distant supervision, so the same label noise that corrupts training also...

arXiv cs.AI

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891...

arXiv cs.LG

A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction

Postoperative acute kidney injury (AKI) after major non-cardiac surgery carries substantial morbidity, yet early intraoperative risk stratification remains difficult. In this retrospective cohort study, we propose SynerT, a waveform-only hybrid temporal backbone that combines a causal dilated TCN with a hierarchy...

arXiv cs.AI

Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order...

arXiv cs.LG

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering,...

arXiv cs.AI

Reinforcement Learning with Decomposed Subtasks

Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental...

arXiv cs.LG

QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs

Relational Deep Learning (RDL) models multi-table databases as heterogeneous temporal graphs, and graph transformers currently achieve state-of-the-art performance on benchmarks like RelBench. However, the current leading model, RelGT, suffers from two key limitations: its random local sampler yields loosely...

arXiv cs.AI

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning...

arXiv cs.LG

Marginally Correct Tool Caches Can Reverse Group-Normalized Policy Updates

Tool-result caching reduces repeated execution in agent training, but also couples rollout randomness. We study a two-action model in which independent and shared execution preserve every rollout's conditional reward distribution. Despite this marginal agreement, sharing one stochastic result per group can reverse...

arXiv cs.AI

Are Stated Reasoning Steps Causally Load-Bearing?

Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure...

arXiv cs.LG

PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation

Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a...

For authors

Submitting to the journal

Please submit your manuscript, in English or German, as either a LaTeX package or a PDF. Every submission is reviewed by the editors, and every author receives a reasoned response.

Charges
Submission is free. A single CHF 450 charge applies when a manuscript proceeds to external peer review.
Review
Manuscripts selected for peer review are assessed double-blind by at least 2 independent specialists outside the editorial team. Peer review is completed within 4 weeks of submission.

Authors

Research from academia and professional practice

The journal publishes work by researchers, doctoral candidates and practitioners in industry and public institutions. Every manuscript is assessed on its method, evidence and contribution.

Named authors take responsibility for the manuscript and record their individual contributions.

Institutions

Collaborative work across laboratories, companies and public bodies

The journal publishes collaborative research conducted across universities, laboratories, companies and public bodies. Evidence from systems in operation is central to its scope.

Named individuals hold authorship; institutions are recorded as affiliations on the published article.

Who decides

Editorial board

Editorial decisions are assigned by subject competence. Each manuscript selected for peer review is assessed double-blind by independent specialists. Reviewer identities remain confidential, and the responsible editor makes the final publication decision.

Editor-in-chief

Dr.Walter KurzMBA, M.Sc.

  • Enterprise AI architecture
  • Multi-agent AI systems
  • AI governance and regulatory compliance
  • AI in company valuation and risk
  • Distributed ledger infrastructure
  • Universität Graz
  • Universität Augsburg
  • FH CAMPUS 02
  • Swissi Institute for AI

Editorial board

Prof. Dr.Velimir Dedić

  • Computer and AI system security
  • Management information systems
  • Data analysis and applied statistics
  • IT security policy
  • E-learning and instructional design
  • FITI Belgrad (Faculty of Information Technology and Engineering)
  • BK University
  • IRRODL

Editorial board

Prof. Dr.Stefan Marx

  • Internal control systems
  • Statutory audit under HGB and IFRS
  • Corporate governance and compliance systems
  • Internal audit and special audits
  • Sustainability and ESG reporting
  • HfWU Nürtingen-Geislingen
  • Steinbeis-Beratungszentrum Corporate Governance und Wirtschaftsprüfung

Editorial board

Prof. Dr.Frank Reinhardt

  • Tax law
  • Banking regulation and supervision
  • Corporate governance and internal control
  • Compliance and risk management
  • HfWU Nürtingen-Geislingen
  • Steinbeis-Beratungszentrum Corporate Governance und Wirtschaftsprüfung

Editorial board

Prof. Dr.Jörg Westphal

  • AI adoption in sales organisations
  • Trust in AI systems
  • AI governance in organisations
  • AI-supported sales training
  • B2B sales management and enablement
  • FOM Hochschule
  • Helmut-Schmidt-Universität Hamburg
  • Universität Hamburg
  • DHBW Stuttgart
  • zfuw Koblenz
  • UCAM Murcia
  • Marketing Management Journal
  • AKAM

Editorial board

Dr.Ralf Kittelberger

  • Compliance management systems
  • Corporate governance and supervisory board duties
  • Employment law and HR compliance
  • AI in legal work and legal operations
  • Company and commercial law
  • HfWU Nürtingen-Geislingen
  • TU Darmstadt
  • Eberhard-Karls-Universität Tübingen
  • Fortbildungsinstitut der Rechtsanwaltskammer Stuttgart

Editorial board

Dr.Florian Kollberg

  • AI risk management in regulated finance
  • Model risk and validation
  • Financial regulation and FINMA requirements
  • ESG reporting and sustainable supply chains
  • Corporate finance, M&A and company valuation
  • HfWU Nürtingen-Geislingen

Editorial board

Dr.Wojtek StrickerM.Sc.

  • AI in company valuation
  • Corporate finance and investment cases
  • KPI and controlling systems
  • Group steering and strategic controlling
  • Renewable energy finance
  • Signum Magnum College
  • FernUniversität in Hagen
  • Frankfurt School of Finance & Management
  • Corporate Finance Institute

Editorial board

Prof. Dr.Svetlana Andjelic

  • Computer adaptive testing
  • Assessment design and measurement
  • Database systems and data modelling
  • Software architecture and domain-driven design
  • Education technology
  • Singidunum University Belgrade
  • Union University
  • FON, University of Belgrade
  • ITS Belgrade
  • Faculty of Computer Science Banja Luka

Editorial board

Prof. Dr.Nenad Dedić

  • Applied cryptography
  • Cloud and data security
  • Software supply-chain security
  • Secure AI architecture and deployment
  • Distributed systems
  • Boston University

Editorial board

Prof. Dr.José Machado

  • Industrial automation and robotics
  • Control systems and mechatronics
  • Industrial systems modelling and simulation
  • Industrial digitalisation and Industry 4.0
  • Manufacturing systems engineering
  • University of Minho (MEtRICs Research Center)
  • École Normale Supérieure de Cachan

Editorial board

Prof. Dr.Šemsudin Plojović

  • Applied statistics
  • Management information systems
  • Knowledge management and e-business
  • Innovation ecosystems and entrepreneurship
  • University of Novi Pazar

Editorial board

Prof. Dr.Ivica Stankovic

  • Quantitative risk modelling
  • Model risk management and validation
  • AI governance in financial institutions
  • Market and credit risk analytics
  • University College Dublin
  • University of Belgrade

Editorial board

Prof. Dr.Enes Sukic

  • Information systems and service architecture
  • Scholarly publishing and peer review
  • Research methodology and reporting
  • Management of technology
  • FITI Belgrad (Univ. Union-Nikola Tesla)
  • University of Niš
  • University of the Balearic Islands

For readers and libraries

Terms of publication

These publication terms apply to every article and form the journal record used by libraries, indexes and funders.

ISSN
3043-1921
Key title
Swissi AI journal
Abbreviated key title
Swissi AI j.
Publisher
Swissi Holding AG, Zug, Switzerland
Editor-in-chief
Dr. Walter Kurz, MBA, M.Sc.
Publication model
Continuous publication. Each article receives an article number.
Review
Manuscripts selected for peer review are assessed double-blind by at least 2 independent specialists from outside the editorial team. Peer review is completed within 4 weeks of submission. Review is double-blind: author and reviewer identities are withheld from each other.
Access
Open access. Full text is available from publication.
Charges
Submission and reading are free. A single CHF 450 charge applies when a manuscript proceeds to external peer review.
Licence
Creative Commons Attribution 4.0 International. Authors keep copyright.
Preservation
Every article is deposited with Zenodo, operated by CERN.
Listed in
ROAD, the ISSN International Centre's directory of open access scholarly resources. The ISSN Portal record is public in full, and the ISSN was assigned by the ISSN Centre Switzerland at the Swiss National Library.
Language
Articles are published in English with a German HTML reading version. The English text is the version of record, and each article PDF is published in English. Manuscripts are accepted in English or German.

Editorial standard

What the editors read for

Every submission is assessed against six criteria, including reproductions and negative results.

MethodReproducible from what is written
The method identifies the data and provenance, model, training and inference settings, and the decisions needed to reproduce the work.
EvaluationBaselines, conditions and failures
The evaluation reports baselines, ablations, operating conditions, failures and the setup behind each central result.
ClaimsConclusions stay within the evidence
A benchmark result is reported as a benchmark result, and a claim about behaviour in service rests on measurements from service.
Prior workWhat it stands on, and what it adds
The article states the literature it stands on and what it adds to it. Where a finding contradicts published work, it engages with that work directly.
Systems in serviceDeployed systems, as they behave
Reports on deployed systems state operating conditions, observed behaviour, measured effects and failures.
Reproduction and negative resultsFull editorial weight under the stated criteria
Reproductions, refutations and negative results carry full editorial weight when the method and evaluation meet the journal's criteria.

Scientific integrity

The journal applies these standards.

Kodex Wissenschaftliche Integrität

Swissuniversities, the Swiss National Science Foundation and Innosuisse drew up a code of conduct for scientific integrity together, under the lead of the Swiss Academies of Arts and Sciences.

PDF (DE)akademien-schweiz.ch

1.00 MB · 4 May 2021

© 2021 Akademien der Wissenschaften Schweiz. Open-access publication under CC BY 4.0, source doi.org/10.5281/zenodo.4707584.

Recent Daily Papers

From Hugging Face

Hugging Face

Hugging Face Daily Papers

Knowledge Pull Requests for Continual Document Authoring

We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or...

Hugging Face

Hugging Face Daily Papers

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate...

Hugging Face

Hugging Face Daily Papers

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long...

Hugging Face

Hugging Face Daily Papers

Calibration as a First-Class Criterion in LLM Evaluation

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks...

Hugging Face

Hugging Face Daily Papers

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and their evaluations, it...

Hugging Face

Hugging Face Daily Papers

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes...

Hugging Face

Hugging Face Daily Papers

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the tt whisper-large-v3-turbo variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced...

Hugging Face

Hugging Face Daily Papers

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and conventional simulators require labor-intensive construction. We present Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned...

Hugging Face

Hugging Face Daily Papers

MemoryAthena: Adaptive Routing over Latent and Generated Memories

Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct Engram retrieval...

Hugging Face

Hugging Face Daily Papers

On the Diffusibility of High-Dimensional Latents

Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction...

Hugging Face

Hugging Face Daily Papers

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal...

Hugging Face

Hugging Face Daily Papers

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and...

Hugging Face

Hugging Face Daily Papers

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial...

Hugging Face

Hugging Face Daily Papers

HappyWorld-Bench

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with...

Hugging Face

Hugging Face Daily Papers

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This...

Hugging Face

Hugging Face Daily Papers

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and...

Hugging Face

Hugging Face Daily Papers

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over...

Hugging Face

Hugging Face Daily Papers

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information...

Hugging Face

Hugging Face Daily Papers

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent...

Hugging Face

Hugging Face Daily Papers

RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling

Reinforcement learning (RL) is vital for optimizing video generation models, with a robust reward model (RM) serving as the cornerstone. However, existing video reward models often produce unstable scalar scores because they directly map complex, subjective video quality into a single score without explicit...

This site stores functional cookies for language, consent and regional routing, plus the theme in browser storage. Choose your cookie setting. Privacy policy