Swissi AI Journal

Open access

AI research for systems that carry responsibility

Articles on autonomous agents, regulated finance, digital identity, ledgers, energy markets, valuation, education and insurance. Read the latest work and follow the ideas across the journal.

SAIJ-5kdnql4rsq27Jul 2026

Multi-Jurisdictional Legal Identity Assurance for Capability Gating

A Design-Science Proposal for Tiered, Reusable Identity Assurance of Natural, Juridical, and Machine Entities

DOI 10.5281/zenodo.21901241

Walter Kurz

Swissi Institute for AI

Flat maximum verification charges every participant for the rarest high-risk case, and excludes those who cannot clear a bar they never needed to. This model holds the assurance state apart from the capability gate that consumes it, so identity demand follows the act and the weight of its consequences rather than mere presence.

Keywords:
  • identity assurance
  • capability gating
  • tiered and reusable verification
  • multi-jurisdictional identity
  • data minimisation
  • entity taxonomy
  • bitemporal reliance
  • design science research

SAIJ-zpd6gvtrfaunJul 2026

Credentials and Triangulated Trust Signals on a Single Accountable Identifier

A Hash-Anchored Distributed-Ledger Framework for Portable Identity across Jurisdictions

DOI 10.5281/zenodo.21901243

Walter Kurz

Swissi Institute for AI

Digital identity stays rigid while it is bound to provider accounts, mutable handles and local wallet schemes. An accountable hash-anchor tier sits above them: an inert root anchor, unlinkable profile anchors for distinct contexts, and gate-specific assurance evaluated at a point in time.

Keywords:
  • digital identity
  • verifiable credentials
  • accountable pseudonymity
  • distributed ledger
  • selective disclosure
  • identity assurance
  • self-sovereign identity
  • hash anchor

SAIJ-sh27g6sykt2kJul 2026

Identity-Staked Consensus and Collusion Resistance in Chartered Validator Sets

A Trust Model for Decentralised and Compliant Distributed Settlement Infrastructure

DOI 10.5281/zenodo.21901245

Walter Kurz

Swissi Institute for AI

Permissioned ledgers are commonly dismissed as centralised because admission is restricted. Separating permissioning from control distribution makes validator identity externally costly collateral: public legal identity, charter state, liability and audit exposure, with affiliation-aware voting caps and per-member collusion margins.

Keywords:
  • identity-staked consensus
  • permissioned ledger
  • proof-of-authority
  • validator trust model
  • collusion resistance
  • settlement infrastructure
  • actor assurance
  • threshold class coverage
  • ledger evidence record

SAIJ-f3jtignfyge3May 2026

Firm Valuation When AI Shapes the Business Model

A Milestone-Based Real-Options Framework for the AI Valuation Uncertainty Problem

DOI 10.5281/zenodo.21901247

arXiv 2609.24181

Walter Kurz1, Wojtek Stricker1, Stefan Marx2, Frank Reinhardt2, Florian Kollberg2

1Swissi Institute for AI2Hochschule für Wirtschaft und Umwelt Nürtingen-Geislingen

Discounted cash flow, the IDW S 1 income approach and market multiples compress milestone probabilities, continuation options and risk shifts into opaque aggregate parameters. A milestone-gated real-options overlay decomposes that value into auditable components, with a Success Readiness Index deriving per-option probabilities from structured pairwise comparisons.

Keywords:
  • firm valuation
  • AI integration
  • real options
  • milestone-based valuation
  • intangible assets
  • AHP
  • multi-criteria decision analysis

SAIJ-cwo7xrcdsautMar 2026

Functional Architecture of European Electricity Trading Markets

Requirements for AI Supported Trading Systems under Regulatory Constraints

DOI 10.5281/zenodo.21901249

Walter Kurz, Wojtek Stricker

Swissi Institute for AI

European electricity trading runs as a constrained multi-layer system in which legal design, exchange microstructure and network physics execute jointly across forward, day-ahead, intraday and balancing horizons. The paper specifies an AI-supported trading architecture with a permission gate on executable actions and fail-closed control logic under REMIT, MiFID II, MiFIR and EMIR.

Keywords:
  • EU electricity market
  • market coupling
  • NEMO topology
  • electricity balancing
  • AI trading systems
  • compliance-by-design

SAIJ-xz3bi3q7fwimAug 2025

Compliant AI Infrastructure for Regulated Finance

A tiered multi-agent framework with DLT audit trails for financial operations in DACH

DOI 10.5281/zenodo.21901251

arXiv 2609.27632

Walter Kurz, Reinhard Magg

Swissi Institute for AI

Regulation is treated as an orientation layer rather than a deterministic ruleset: a matrix of regulatory intent and exposure is compiled into concrete prohibitions, obligations and runtime budgets. Evidence, decisions and reason codes bind to a permissioned DAG, so a supervisor can replay how an outcome was reached and attribute failure.

Keywords:
  • DACH finance
  • regulated financial institutions
  • multi agent expert system
  • policy compiled orchestration
  • objective under constraints
  • permissioned DLT
  • DAG timestamping
  • audit trails
  • EU AI Act
  • MiFID II
  • DORA
  • GDPR
  • human oversight
  • execution gating
  • ESG budgets
  • verification and assurance

SAIJ-ddkjais6s332Aug 2025

A regulatory-compliant AI and verification system for higher education under ESG-aligned constraints

DOI 10.5281/zenodo.21901253

Walter Kurz, Michel Malara, Wojtek Stricker

Swissi Institute for AI

Two linked components for higher education: a role-specific multi-agent framework for institutional operations, and a decentralised verification layer for audit, credential authentication and tamper-evident records. GDPR, the EU AI Act, EQF, ECTS and ESG directives are encoded as structural constraints rather than checked after the fact.

Keywords:
  • Regulatory technology
  • artificial intelligence in education
  • multi-agent AI systems
  • decentralised verification
  • academic tokenisation
  • GDPR compliance
  • EU AI Act
  • digital credential infrastructure
  • ESG governance
  • UniAI
  • UniDVS

SAIJ-zkihhbpahsbrAug 2025

Verifiable Federated AI Infrastructure

Swiss compliant federated AI DLT network using Nash equilibrium and ESG metrics

DOI 10.5281/zenodo.21901255

Walter Kurz, Michel Malara, Velimir Dedić

Swissi Institute for AI

Centralised AI infrastructure scales, and collides with latency, auditability and energy constraints. The design separates centralised training from decentralised inference and storage across five node classes, tying a size-neutral availability floor to tiered rewards for service level, ESG performance and anti-concentration.

Keywords:
  • decentralised data centre
  • AI
  • Federated AI infrastructure
  • ESG
  • ESG-aware compute
  • Nash equilibrium
  • digital sovereignty
  • Swiss data regulation
  • tokenised infrastructure
  • verifiable AI services

SAIJ-soeptiqyucowAug 2025

Generic AI-DLT Enterprise System

Architecture and methodology for scalable domain adaptation from a unified core framework

DOI 10.5281/zenodo.21901257

Walter Kurz, Michel Malara, Velimir Dedić

Swissi Institute for AI

Compliance in AI deployments is usually applied afterwards, through prompt engineering, rather than built into the foundation. This architecture encodes regulatory, governance and ESG requirements as an objective-under-constraints problem, so every specialised agent operates within legally admissible and auditable bounds before any domain work begins.

Keywords:
  • Compliance-first AI
  • Multi-agent systems
  • Distributed ledger technology
  • Directed acyclic graph
  • Regulation by design
  • ESG integration
  • Domain-agnostic architecture
  • Deployment-agnostic architecture
  • Vendor-agnostic architecture
  • Objective-under-constraints

SAIJ-qzvrl4bwy7y2May 2025

Multi-Agent AI Architecture for Regulated Insurers

A generic AI framework under Solvency II and the AI Act in Austria and Germany

DOI 10.5281/zenodo.21901259

arXiv 2609.27636

Walter Kurz

Swissi Institute for AI

The insurer is modelled as a constrained optimisation entity under solvency, legal, ESG and operational boundaries, then decomposed into specialised agents for capital, underwriting, claims, compliance and fraud. Human-in-the-loop roles enter through tiered access control, with an orchestrator enforcing regulatory admissibility across the set.

Keywords:
  • Multi-Agent Systems
  • Enterprise AI
  • Insurance Firms
  • Solvency II
  • AI Act
  • Regulated Environments
  • Constrained Optimisation
  • Principal-Agent Theory
  • Nash Equilibrium
  • Arrow’s Risk Pooling
  • Austria
  • Germany
  • Institutional Design
  • ESG Compliance
  • Regulatory Architecture
  • Model Context Protocol (MCP)
  • Agent-to-Agent Protocol (A2A)
  • AI Governance
  • Algorithmic Accountability
  • Financial Regulation

From arXiv

Recent arXiv Articles

arXiv cs.AI

Silent Failures in Agent-Tool Interaction: An Audit of ToolUniverse

Agentic AI systems are increasingly adopting automated pipelines that integrate multiple tools. While prior research and benchmarks have studied about task success and task completion of these agentic systems, the research about agent to tool interaction, specifically in biology agentic workflow is limited. This...

arXiv cs.LG

The Drift Contract: Spectral Updates for Depth-Robust Local Learning

Local learning trains each layer with its own auxiliary loss and no global backward pass, which makes layer updates structurally parallel. Two problems have kept it marginal: accuracy degrades as depth grows, and hyperparameters are fragile. We apply Muon-style spectral update geometry (momentum orthogonalization...

arXiv cs.AI

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Modern language-model agents are built around the agent loop, where the LLM is placed in an environment exposing a set of tools, and the LLM has full control over the workflow by alternating between tool calls and observing their output. However, certain workflows currently require additional engineering beyond the...

arXiv cs.LG

Signal2Symbol: Neuro-Symbolic Temporal Reasoning for Explainable Physiological Time-Series Anomaly Detection

Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit complex temporal structure, substantial acquisition variability, and a strong need for transparent decision-making. Although deep models can achieve high detection performance, they often provide limited insight into...

arXiv cs.AI

TwinCheck: Evidence-Grounded Negative-Twin Verification for Stateful Tool Agents

A single locally plausible tool call can derail an otherwise successful agent trajectory. Suspicion alone does not justify intervention, because the replacement itself can introduce the very failure verification is meant to prevent. We introduce TwinCheck, an inference-time verification policy that considers...

arXiv cs.LG

HARN: Hierarchical Associative Resonance Network for Event-Driven Multi-Timeframe Forecasting

Financial time series evolve across multiple temporal resolutions, challenging forecasting systems to incorporate newly available information without repeatedly recomputing unchanged representations. We introduce HARN, a Hierarchical Associative Resonance Network for event-driven multi-timeframe forecasting. HARN...

arXiv cs.AI

Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations

The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational...

arXiv cs.LG

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be...

arXiv cs.AI

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a...

arXiv cs.LG

LWCal: Loss-Weighted Calibration for Tabular Classifiers with Noisy Calibration Labels

Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deployment settings, however, labels come from weak annotators, historical decisions, heuristics, or distant supervision, so the same label noise that corrupts training also...

arXiv cs.AI

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

Dependency conflicts in Python ecosystems arise from incompatible version constraints, missing packages, and undocumented compatibility relationships, causing many real-world code snippets to fail at execution. This paper presents PLLM+, a hybrid dependency-repair pipeline evaluated on the HG2.9K benchmark of 2,891...

arXiv cs.LG

A Leakage-Aware Multimodal Evaluation Framework for Early Intraoperative Acute Kidney Injury Prediction

Postoperative acute kidney injury (AKI) after major non-cardiac surgery carries substantial morbidity, yet early intraoperative risk stratification remains difficult. In this retrospective cohort study, we propose SynerT, a waveform-only hybrid temporal backbone that combines a causal dilated TCN with a hierarchy...

arXiv cs.AI

Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order...

arXiv cs.LG

COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation

While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering,...

arXiv cs.AI

Reinforcement Learning with Decomposed Subtasks

Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental...

arXiv cs.LG

QUARTET: Quad-branch cross-Attention and Random-walk Traces for Enhancing Transformers on Relational Graphs

Relational Deep Learning (RDL) models multi-table databases as heterogeneous temporal graphs, and graph transformers currently achieve state-of-the-art performance on benchmarks like RelBench. However, the current leading model, RelGT, suffers from two key limitations: its random local sampler yields loosely...

arXiv cs.AI

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning...

arXiv cs.LG

Marginally Correct Tool Caches Can Reverse Group-Normalized Policy Updates

Tool-result caching reduces repeated execution in agent training, but also couples rollout randomness. We study a two-action model in which independent and shared execution preserve every rollout's conditional reward distribution. Despite this marginal agreement, sharing one stochastic result per group can reverse...

arXiv cs.AI

Are Stated Reasoning Steps Causally Load-Bearing?

Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure...

arXiv cs.LG

PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation

Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a...

Recent Daily Papers

From Hugging Face

Hugging Face

Hugging Face Daily Papers

The Linear Representation Hypothesis Needs a Group Action

To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different...

Hugging Face

Hugging Face Daily Papers

Knowledge Pull Requests for Continual Document Authoring

We introduce Knowledge Pull Requests (KPRs), a framework for continual document authoring that makes each change interpretable. Documents require ongoing revision as new knowledge surfaces from other sources, languages, or times, but existing approaches either edit with no account of what knowledge changed or...

Hugging Face

Hugging Face Daily Papers

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate...

Hugging Face

Hugging Face Daily Papers

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long...

Hugging Face

Hugging Face Daily Papers

Calibration as a First-Class Criterion in LLM Evaluation

Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks...

Hugging Face

Hugging Face Daily Papers

FLEET: From Logits Entropy to Enhanced Trajectories in Text Generation

Solutions based on large language models (LLMs) often rely on temperature sampling to improve accuracy and stability by aggregating multiple samples from the completion distribution. However, this memoryless approach is inherently suboptimal: because it lacks awareness of prior generations and their evaluations, it...

Hugging Face

Hugging Face Daily Papers

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes...

Hugging Face

Hugging Face Daily Papers

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the tt whisper-large-v3-turbo variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced...

Hugging Face

Hugging Face Daily Papers

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI

Scalable simulation is essential for robot data generation, policy training, evaluation, and safe iteration, yet real-world interaction is costly and conventional simulators require labor-intensive construction. We present Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned...

Hugging Face

Hugging Face Daily Papers

MemoryAthena: Adaptive Routing over Latent and Generated Memories

Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. We study whether useful memory can also be generated rather than only retrieved. MemoryAthena uses three pathways: direct Engram retrieval...

Hugging Face

Hugging Face Daily Papers

On the Diffusibility of High-Dimensional Latents

Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction...

Hugging Face

Hugging Face Daily Papers

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

Video is a rich representation of a physical event, capturing appearance, geometry, motion, and temporal evolution. Other modalities, such as 3D body motion or audio, encode narrower aspects of the same event. We find that joint multimodal diffusion transformers exhibit a corresponding asymmetry in cross-modal...

Hugging Face

Hugging Face Daily Papers

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and...

Hugging Face

Hugging Face Daily Papers

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial...

Hugging Face

Hugging Face Daily Papers

HappyWorld-Bench

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with...

Hugging Face

Hugging Face Daily Papers

Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This...

Hugging Face

Hugging Face Daily Papers

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and...

Hugging Face

Hugging Face Daily Papers

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over...

Hugging Face

Hugging Face Daily Papers

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information...

Hugging Face

Hugging Face Daily Papers

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distinguish who said what, whom each statement concerns, how individuals perceive one another, what information is shared by the group, and how states change over time. Recent...

This site stores functional cookies for language, consent and regional routing, plus the theme in browser storage. Choose your cookie setting. Privacy policy