Abstract
In private wealth management, a manager delegating to artificial intelligence (AI) acts as the client's agent and the system's principal. We introduce a model-independent formulation that combines nested principal–agent delegation with constrained joint maximisation as the task assigned to the AI system. The objective represents client and manager outcomes separately over portfolio–workflow pairs. Legal duties, mandate requirements and evidence sufficiency determine admissibility, with Switzerland, Germany and Austria supplying the legal context. Weights and reference-service floors make the trade-off explicit; concession accounting separates their effects on the client. Analytical constructions and a simulation using public-market observations illustrate the approach. Across eight decision states from four constructed mandates, omitted client liabilities caused two liquidity violations, omitted manager terms caused two capacity violations, and mistranslated weights changed four otherwise admissible choices under faithful optimisation. At the declared weights, six states selected a higher service tier than the client-best alternative, with client concessions of EUR 1,178 to EUR 2,264 and manager gains of EUR 3,062 to EUR 10,381. Three instruction forms each reached all 32 specified decisions under shared numerical, evidence and simulated approval controls; professional instructions matched explicit nested delegation on accuracy and clarification count. Subsequent 2022 exchange-rate and yield paths, combined with constructed growth scenarios, produced lower client outcomes than the reference service although the selected services met the decision-time forecast benchmarks. These examples suggest that the approach could help make mandate choices and their consequences easier to examine. Professional and field studies could assess whether this improves oversight and client outcomes.
Introduction
A wealth manager delegating work to a system of AI agents asks it to make decisions from information filtered through two parties. The client may leave circumstances or preferences undisclosed; the manager may incompletely convey that account, its own objectives or the service conditions. Either omission can reflect strategic disclosure, uncertainty or imperfect articulation. The AI can have the least direct access to the private information on which its task depends, while being expected to serve both parties. The specification risk is that gaps are silently filled with learned defaults or unsupported assumptions about their circumstances and objectives, then carried into downstream decisions. Studies of multi-agent failures document action under incorrect assumptions, failures to seek clarification and incomplete verification (Cemri et al. 2025).
Agency theory addresses delegated authority under information asymmetry and potentially divergent interests (Jensen and Meckling 1976). A classical formulation advances the principal’s interests while securing the agent’s participation and desired behaviour (Holmstrom 1979). Financial applications examine moral hazard in delegated portfolio management and incentives behind unsuitable sales (Stoughton 1993; Inderst and Ottaviani 2009). Hierarchical agency supplies a basis for examining successive delegations (Tirole 1986). We propose to apply this framework to the client–manager–AI relationship: the client is the manager’s principal, and the manager is both the client’s agent and the AI system’s principal. This structure locates responsibility for representing interests and transmitting information at each interface.
We propose to specify the AI’s task through two explicit outcome functions: one for the client and one for the wealth manager. Client outcomes include return, liquidity and financial security; manager outcomes include service quality, resource use and commercial value. Their interests can coincide or conflict. The manager’s interests shape the service whether or not the specification names them; naming them turns the trade-off into something a reviewer can read. Constrained joint maximisation over portfolio–workflow pairs makes the trade-off explicit, with legal duties, mandate requirements and evidence sufficiency determining admissibility. Weights and reference-service floors require justification and authorization. The intended engineering use is to translate this specification into agent instructions and workflow rules that preserve both evaluations, bind decisions to supported premises and identify when unresolved information or objective choices require human clarification. Each agent’s assigned role must remain consistent with that shared specification.
The DACH setting makes the treatment of incomplete information consequential. In Switzerland, Articles 11–12 of the Financial Services Act (FinSA) require the provider to seek service-relevant client information. If that information is insufficient to assess appropriateness or suitability, Article 14(1) requires prior notice to the client that the provider cannot perform the assessment (Swiss Confederation 2018, arts. 11–14). For investment advice and portfolio management under EU suitability rules, Article 25(2) of the Markets in Financial Instruments Directive (MiFID II) requires necessary client information, and Article 54(8) of Delegated Regulation 2017/565 bars recommendations when that information is missing (European Parliament and Council of the European Union 2014, art. 25(2); European Commission 2017, art. 54(8)).
The manager must specify both which information the service requires and what an unresolved gap permits the system to do. A uniform hold policy adds an institutional restriction to the Swiss notification rule, while the EU suitability rule expressly restricts recommendations. Applicable mandate and other legal duties continue to govern the service. Jurisdiction and service type enter the decision specification because the same information gap can require different responses.
The research question is how nested client–manager–AI delegation can be translated into an inspectable specification for joint optimisation and clarification under incomplete information. The intended users are the manager responsible for the mandate and the engineer configuring the agents and workflow. The assessment distinguishes the adequacy of a transmitted specification, compliance with it and outcomes under subsequent market conditions. The simulation and executable constructions illustrate how the theory can be instantiated; the system architecture remains the builder’s choice.
Contribution
The contribution is a nested principal–agent formulation in which the delegated AI maximises a joint client–manager objective subject to binding constraints. The manager acts as the client’s agent when agreeing the mandate and as the AI’s principal when specifying its task. We connect these roles through separate client and manager outcome functions over portfolio–workflow pairs, with legal, mandate and evidence requirements determining admissibility before optimisation (5). The weight places each decision between a client-first and a manager-first programme (Proposition 2). Concession accounting separates benchmark restrictions from weighting; a proposed review screen requires policy review and assesses the total concession. Disclosure, translation and execution provide a decomposition of departures under stated counterfactual and tie-selection conventions. A simulation illustrates required-input recovery, numerical selection and version-bound approval. Exact constructions assess the review screen, while partition-based discretionary clarification is implemented for the restricted client selector.
Three research questions guide the assessment. First, how can client–manager–AI delegation be represented by a joint decision objective with binding client-protection, mandate and evidence constraints? Second, how do information omissions, representation choices and uncertainty assumptions affect admissibility, selection and the separate outcomes of both parties? Third, how can the specification be instantiated in agent instructions and workflow controls?
The formal specification and its finite client-selection instance address the first question. Analytical constructions and the portfolio–workflow simulation address the second, varying what each delegation transmits while retaining the original mandate for evaluation. The instruction comparison addresses the third, complemented by deterministic, single-agent and multi-agent implementations. Together these constructions demonstrate how declared objectives, service conditions and evidence requirements govern a decision.
Methods
We used design science to develop the decision specification (Hevner et al. 2004; Peffers et al. 2007), combining analytical constructions, economic simulation and worked implementations for artificial evaluation (Venable et al. 2016). The formal results depend on the stated objectives, constraints and information structure; any executor preserving these rules can instantiate the specification. Exact-rational constructions isolated objective representation, benchmark eligibility, clarification and evidence transitions. Case policies supplied the assessment constraints informed by the legal mapping.
The economic simulation represented complete client circumstances, manager service terms and a declared mandate, then controlled what disclosure and specification transmitted. Four constructed families covered business-sale liquidity, retirement drawdown, currency liabilities and concentrated holdings. Three portfolios and three service workflows competed with a reference service. Separate interventions omitted client liabilities, omitted manager terms or changed the transmitted weight; faithful optimisation of each altered specification was evaluated against the original mandate. Simulated approval checked current premise support, a matching calculation, numerical constraints, reference floors and evidence version. The comparison held substantive mandate authorization fixed; the concession-review screen was assessed separately through exact constructions. Thus numerical execution and domain authorization have distinct roles in the demonstration.
Monthly cash flows included withdrawals, liabilities, currency holdings, zero-coupon bond valuation, fees, forced-sale costs and supervision. Client value combined expected wealth, a downside penalty and declared service value; manager value combined net fee income and service quality. Their scales and reference floors were fixed before selection. European Central Bank exchange rates and government yields supplied 144 monthly observations for 2014–2025 (European Central Bank, n.d.-b, n.d.-a). Growth returns, preferences, service conditions and three scenario probabilities were authored. Decision-time evaluations used the initial market observation; subsequent observed exchange rates and yields entered the outcome calculation with the same constructed growth scenarios.
Professional prose, explicit joint-objective instructions and explicit nested-delegation instructions received the same declared objective, evidence, questions, calculators, resource ceilings and simulated approval controls. Calculators returned all candidate outcomes for agent selection. December 2016 mandates supported development; the same four families at December 2021, with rescaled liabilities, spending and hourly costs, supplied four evaluation cases with eight information conditions each. These covered complete evidence, client or manager gaps, both gaps, translation conflict, refusal, correction after approval and equal-outcome controls. One run per treatment and condition yielded 96 assignments. All three instruction forms used the same model and runtime settings. Private states, answers and subsequent market paths stayed with the evaluator.
The retained studies tested supplied client-selection policies using deterministic, single-agent and multi-agent workflows. The structured pilot contained 108 assignments; each narrative comparison contained 126. Swiss Banking Ombudsman cases supplied authorization, retirement-depletion and investment-cash-flow mechanisms (Swiss Banking Ombudsman 2023, 2020, 2019); records and exposures were authored. The revised narrative interface combined typed extraction, clearer instructions and reusable arithmetic receipts, with re-evaluation on the same cases. These implementations illustrate the same decision concepts through alternative workflow arrangements.
Scoring distinguished fidelity to the original mandate, raw proposal defects, issuance, clarification burden, the declared joint optimum and separate economic outcomes. Current evidence and approval governed issuance. Independent cash-flow identities checked financial arithmetic; saved actions were replayed to verify decisions and evidence transitions. Source and model-response records retained configurations, prompts and usage. Conditions sharing a mandate were treated as dependent observations. Model expenditure and the constructed service costs were reported separately. The accompanying Supplementary Analytical and Resource Results contains the detailed counterexamples, reference traces and resource tables.
The legal mapping concerns investment advice and portfolio management for private or retail clients, with transaction-specific advice and execution services supplying applicability contrasts.
Under Swiss FinSA, transaction-specific advice assesses knowledge and experience for the proposed instrument; portfolio-based advice and management also assess financial situation and objectives. Insufficient information triggers Article 14’s client-information duty (Swiss Confederation 2018, arts. 11–14). German and Austrian advice and portfolio management use the suitability inputs in WpHG section 64(3), WAG 2018 section 56 and Article 54 of Delegated Regulation 2017/565, including loss-bearing capacity and risk tolerance (Federal Republic of Germany 1994; Republic of Austria 2018a; European Commission 2017). Article 54(8) conditions recommendations on obtaining the necessary information (European Commission 2017, art. 54(8)). The artifact’s uniform hold policy is a conservative design choice where the mapped Swiss provision specifies an information duty. The constructed liability amount is classified as required by the case policy.
Swiss execution or transmission services engage Article 13’s notification rule; institutional-client treatment and professional-client waivers require the separate scope of Article 20 (Swiss Confederation 2018, arts. 13 and 20). EU appropriateness and qualifying execution-only services follow the conditions in MiFID II Article 25(3)–(4), with professional-client assumptions addressed separately in the delegated rules (European Parliament and Council of the European Union 2014, art. 25; European Commission 2017, arts. 54–56). GDPR Articles 5, 6 and 9 govern processing principles and relevant grounds; Swiss processing is assessed under the Federal Act on Data Protection, including Articles 6 and 30–31 (European Parliament and Council of the European Union 2016; Swiss Confederation 2020). Client willingness, evidential adequacy and lawful processing receive separate policy decisions.
The legal analysis uses a cutoff of 30 September 2026: the FinSA consolidation of 1 March 2024, Regulation 2017/565 consolidation of 2 August 2022, current German section 64 text and Austrian section 56 effective 9 June 2022. MiFID II Article 25 follows the consolidation of 6 June 2026. The German section 64 text is the version in force at the cutoff, cited without a provision-level effective date (Swiss Confederation 2018; European Commission 2017; Federal Republic of Germany 1994; Republic of Austria 2018a; European Parliament and Council of the European Union 2014). FINMA Circular 2025/2, effective 1 January 2025, addresses service classification and knowledge-and-experience elicitation for providers within its scope (Swiss Financial Market Supervisory Authority 2024, paras. 2, 4 and 13–14).
The constructed legal contrasts concern domestic retail services. Cross-border application depends on provider authorization, client category and the origin of the service request. Swiss territorial rules and German or Austrian market-access requirements can apply to the same relationship (Swiss Federal Council 2019, art. 2; Federal Republic of Germany 2021, sec. 15; 1961, sec. 32; Republic of Austria 2018b, secs. 21, 23–24). Exclusive client initiative has a fact-dependent scope, assessed from the service and solicitation history (European Securities and Markets Authority 2021). Market-access conditions, suitability requirements and the territorial scope of data-processing rules are assessed separately (European Parliament and Council of the European Union 2016, art. 3; Swiss Confederation 2020, art. 3).
From Classical Agency to Nested Delegation
Agency theory starts from a principal who cannot observe what the agent does or knows (Jensen and Meckling 1976). In the hidden-action formulation, the principal chooses a payment schedule and the effort it wants the agent to supply (Holmstrom 1979, 75–76, eqs. (1)–(3)):
Three features carry this model. The agent has interests of its own, expressed in and . The agent holds the informational advantage, since it alone observes . And the principal’s instrument is the contract: it steers the agent through what it pays.
Wealth management with AI contains two delegations. The client delegates advisory work to the wealth manager, and the manager commissions and supervises the AI system:
1 shows the manager translating the client’s mandate into AI instructions and receiving proposed advice with supporting premises. The dashed link denotes permitted client-context elicitation.
Note. The wealth manager acts as the client’s agent () and the AI system’s principal (). Solid arrows show the two delegation relationships; the dashed connection shows permitted client-context elicitation. Source: the authors’ conceptual model.
The instrument changes. We model the AI as an executor of the supplied specification, with no separate utility, participation decision or strategic choice of effort. Computational and supervisory costs enter the manager’s service evaluation. Under this assumption, the participation and incentive constraints of Equation 1 apply to the human service relationship; the second delegation is governed by the specification. The manager steers the system through the specification it hands down. contains an objective , a set of alternatives the system may recommend, and rules for the information it may rely on. Where the classical agent maximises its own utility under the contract, the system is asked to maximise the objective it was given, on the input it received:
Information passes through two interfaces. In Equation 1 the agent knows more than the principal. In the nested chain, the AI receives the client’s private circumstances and the manager’s service terms through disclosure and specification. Its access to that information depends on what those interfaces transmit and what supplementary evidence it can obtain. Let describe the client’s circumstances and preferences and the manager’s service objectives, costs and incentives. With a recorded mandate and jurisdiction and service conditions , information reaches the AI in three steps:
Disclosure can itself be a choice. A client whose true utility differs from the declared evaluation, for reasons of privacy, tax or attachment to a holding, selects what to disclose with the resulting recommendation in view:
Proposition 1 (indistinguishable client states). Fix the mandate, policy primitives and decision rule. Let two client states induce the same complete AI decision input , but have disjoint sets and of fully informed client-optimal portfolios under those primitives. A rule selecting a portfolio solely from can guarantee a fully informed client optimum in at most one of these states.
A deterministic rule receives the same input and selects the same portfolio in both states. That portfolio cannot belong to both disjoint optimal sets. A randomised rule has the same conditional output distribution in both states; probability-one optimality in each would require its support to lie in their empty intersection.
The result concerns selection at the given input. A permitted question can distinguish the states; a hold preserves the unresolved decision. The proposition identifies an informational limit independently of the system’s computational ability.
The conflict receives an explicit evaluation. The manager is agent in the first delegation and principal in the second. As agent it owes the client loyalty; as principal it writes the system’s objective. The proposed specification records the weight and scale assigned to its commercial interests together with the alternatives and reference service. Existing governance already requires documentation and review of automated decision rules (European Securities and Markets Authority 2022, para. 90). Explicit dual evaluations add a measurable trade-off and a comparator against which a reviewer can examine each client concession.
We take the service arrangement and its remuneration as given. The mandate is also incomplete: circumstances change, so elicitation and specification have to be renewed (Grossman and Hart 1986; Hart and Moore 1990). How client and manager respond strategically to a specification that both can read is a further layer of analysis, taken up in 9.
The system decides on a pair : a portfolio recommendation and the workflow that delivers it, drawn from a finite set . Let contain the accepted, currently valid case evidence from . The declared evaluations and represent client and manager outcomes. Client criteria may include return, liquidity and security; manager criteria may include service quality, delivery cost and supervisory effort. Their definitions, scales and treatment of uncertainty are fixed before any comparison and remain subject to mandate review. An expectation needs an explicit belief model, a robust evaluation needs declared plausible states, and probabilities generated by a model need separate calibration before they serve as beliefs about a particular client.
Let index the legal, risk, processing, mandate and independently justified service-viability constraints, including applicable client-interest duties, with residuals when satisfied. Service viability records necessary delivery resources; preserving a reference profit is a separate benchmark choice. Let indicate that the information necessary for the service is established, and let be the non-empty set of states compatible with the evidence. Equation 7 admits the alternatives that pass the required-information gate and satisfy every constraint in every such state:
A declared reference service gives each party a floor , , under the same evidence and evaluation rules. The alternatives that are admissible and keep both parties at their reference level form the feasible set
Within the first delegation can be written in the form of Equation 1. A client who could choose the service itself would solve
What should a manager who takes both of its roles seriously write into ? Programme Eq. W treats the second delegation as if the first did not exist. Programme Eq. C looks like the loyal choice, and it has a weakness of its own. The workflow uses the manager’s staff time and supervisory capacity and determines the fee the client pays, so the manager’s interests are part of the decision whether or not the objective names them. Under a client-only objective they enter elsewhere: in which alternatives are offered, which service tiers exist and how much review a case receives. Those choices are made outside the specification, where a reviewer sees no trade-off because none is written down.
We therefore propose that the specification states both evaluations and the weight between them. For a fixed , Equation 9 selects
Proposition 2 (limiting programmes). Let be non-empty. Every solution of Equation 9 is Pareto-efficient in with respect to the two declared evaluations. There are thresholds and such that every solution solves Eq. W when and solves Eq. C when .
An alternative in that weakly improved both evaluations of a solution and strictly improved one would have a strictly larger weighted sum, which contradicts optimality. For the thresholds, let be the smallest gap between the highest client value in and any lower client value, and let be the range of manager values in . An alternative below the highest client value loses at least and gains at most against a client-best alternative, so it cannot be selected when . The argument for exchanges the two parties. If all alternatives share the highest client value, every solution solves Eq. C for any .
The efficiency statement runs in one direction. On a finite set an efficient compromise can be selected at no weight: with client and manager values , and , the third pair scores at every while one of the first two scores at least . A mandate that wants such a compromise states it as a constraint.
The weight therefore has a plain reading: it places the decision between the manager’s programme and the client’s programme, and client priority is the case . The proposal leaves the value of to the mandate. It requires that the value be declared, justified against the mandate and authorized before the system acts on it.
The other elements have a reading in the nested structure as well. Admissibility comes first because the first delegation ranks above the second: legal duties and the mandate bound what the manager may want, so manager value competes only among alternatives that the client relationship already permits. The floors tie the decision to the reference service. The reservation utility in Equation 1 is the value of an outside option; the manager floor is stricter, since it preserves the full managerial value of the reference service, and this choice needs its own justification in the mandate. An empty calls for revised alternatives or a held decision, and current human approval remains a condition of issuance.
Review requires the cost of both benchmark restrictions and weighting. Let solve Eq. C. Relative to that floor-feasible client optimum, the weight concession and manager gain are
Delivery requirements enter . New evidence of a binding requirement triggers a revised admissible set and recomputation of , and . This preserves comparison with the best remaining deliverable service. Let record completed independent policy review of the objectives, scales, weights, reference floors, service viability and applicable duties. Let be a supported, current incremental client benefit omitted from the declared evaluation, measured on its client scale; benefits already represented receive zero additional credit. We propose the following institutional screen for submitting a decision to substantive authorization:
The chain in Equation 5 supplies a counterfactual decomposition of departures from an intended decision. Let be the weighted sum of Equation 9 under the authorized mandate and complete information . Fix a common deterministic rule for selecting among tied maximisers. Let be the choice made under those conditions, the choice the same rule makes from what the client disclosed, the choice it makes from the specification the manager transmitted, and the recommendation the system issues. For stages with a selected alternative, then
A system that selects the transmitted optimum using the same tie rule makes and drives the execution term to zero. Membership in the argmax of Equation 4 alone permits different tied choices. For example, transmitting makes and tie in the construction in 3, while the authorized weight gives them values and . Choosing and faithfully issuing assigns to execution under . If the executed tie rule is unspecified, report the range of over transmitted maximisers; this example gives . Regret measured under the transmitted objective is zero for every faithful maximiser. Responsibility for the specification remains with the manager under either reporting convention.
The specification fixes five things: two outcome functions that stay separate through every calculation, an authorized trade-off between them, admissibility checked before any comparison, an owner for every premise, and approval bound to the current evidence and policy version. The client confirms preferences, financial objectives and agreed service terms. The manager proposes delivery costs, commercial criteria, scales, the trade-off weight and a reference service. A domain reviewer independent of that commercial choice assesses the policy and each decision’s compatibility with the client mandate and applicable duties, including both benchmark and weight concessions in . Domain review also establishes the service and evidence requirements that permission checks apply before comparison and issuance. The responsible manager authorizes the reviewed policy version. Client agreement records preferences and service scope; applicable legal duties remain binding constraints (European Securities and Markets Authority, n.d., para. 1 and 10).
In the restricted client selector, the manager specifies candidates, facet utilities, references and clarification thresholds; calculation tools apply Pareto filtering and assess the declared answer partitions. The agent identifies the premises supporting its proposal, while the evidence service tracks their validity and source ancestry. Human approval binds the candidate to the current evidence and policy version.
These responsibilities can be implemented through a single agent, several agents or a deterministic procedure. The analytical constructions and simulations in 6 examine the corresponding decision and evidence rules. Conformance establishes fidelity to the authorized specification; mandate and domain review assess its adequacy. A material change in objectives, alternatives, evidence criteria or permitted actions requires policy review and renewed assessment.
The finite client selector instantiates Equation 9 when the workflow and manager value are fixed and the floors exclude no feasible candidate. At an evidence state satisfying , let admissibility match and client value be a strictly increasing normalisation of the robust Nash product. The maximisers then coincide with Equation 18; restricting candidates to gives the Pareto selector. 2 and Algorithm 1 specify clarification and approval for this restricted case. The facets represent dimensions within the client outcome function. Discretionary clarification for the full joint objective, including a contingent sequence of questions, requires a further model of both parties’ outcomes and answer probabilities or a declared robust question policy.
The evidence representation adapts the assertion/source distinction and reliance-event model developed for legal identity assurance (Kurz 2026). Each premise carries its origin, supporting evidence, period of validity and use in the recommendation. Evidence sufficiency depends on the premise: a declared preference can be relevant on the client’s authority, while a claim about an external liability may require corroboration.
Within the restricted client-selection procedure, facet agents represent client objectives through manager-approved utilities, reference values and evidence dependencies. The shared policy fixes their relation to the joint specification; agent decomposition must preserve that representation. Organisational roles supply hard legal, risk and data constraints. Evidence records retain source ancestry, so outputs derived from one premise share an origin and its subsequent invalidation (Kurz 2026). Accepted factual evidence filters the enumerated joint-state set while preserving its declared dependencies.
Each context gap has a purpose, a decision relevance assessment and a permitted resolution route. The evidence officer assesses an answer’s evidential adequacy and processing authorization. Acceptance filters the state set; refusal preserves the gap and its applicable route. Client willingness, evidential sufficiency and lawful processing remain separate decisions, including for special-category and third-party data (European Parliament and Council of the European Union 2016, arts. 5–7, 9 and 14).
Let be a finite candidate set and the accepted, currently valid evidence. The policy supplies a finite joint-state set; compatible evidence selects the non-empty subset . A conflict that leaves this subset empty holds the decision for resolution. For each active facet , Equation 16 takes the lowest utility across the evidence-compatible states:
In the evaluated procedure, fixed facets represent client objectives. Equation 18 defines the provisional robust maximiser set using the client Nash-product rule:
The selection rule first removes candidates dominated across the current state/facet utility entries. Write for those for which no satisfies for every , with strict inequality for at least one entry. Equation 19 applies the Nash-product comparison to this statewise-undominated set:
Let index the unresolved inputs in the fixed declared schema whose possible recorded values induce nontrivial partitions. For input , partitions by its possible recorded values. For each nonempty answer cell , calculate , and using the same primitives. With nonempty current feasibility, every current candidate remains feasible in each cell. For a current maximiser , Equations 20 and Eq. 21 define separate monetary and client-criterion consequences of changing the selection after one declared input is resolved:
The policy declares thresholds . Input blocks candidate when either consequence is positive and reaches its corresponding threshold. Equation 23 retains current maximisers whose consequences remain zero or below the corresponding threshold for every unresolved input:
The responsible manager, supported by domain review, must authorize both thresholds and their materiality rationale before execution. Monetary materiality is specified in euros for the mandate; client-criterion materiality requires comparisons of preference consequences on the declared utility scales. Zero thresholds block every positive consequence. The constructed checks stipulate threshold values to examine policy behaviour; practical calibration remains prospective.
Mandatory information, legal and provenance conditions govern every candidate. After these gates pass, a nonempty permits review of its members. If feasibility is nonempty but every current maximiser is blocked, the scheduler asks the first permitted unresolved input that blocks at least one current maximiser, within the fixed schema order and budget. It assesses declared partitions even when a question is unavailable or refused; those restrictions govern the lawful resolution route. Mandatory refusal, exhausted resources or unavailable consequential questions preserve a hold. A new accepted answer triggers reassessment, and a service-scope revision triggers a fresh applicability assessment. Necessary information remains a precondition for the relevant EU advisory service after human escalation (European Securities and Markets Authority 2016).
For comparison, we evaluate a global monetary clarification rule using the unfiltered and singleton-state feasible and winner sets . Equations 24 and Eq. 25 define its global change indicator and monetary-consequence measure:
2 connects the selection and clarification rules to the recommendation workflow. Mandatory checks precede comparison, consequential gaps prompt clarification or referral, and approval binds the selected candidate to the current evidence and policy version.
Note. is the set of review-eligible maximisers at evidence state . Approval binds a candidate to the current evidence and policy version. Source: the authors’ recommendation procedure.
Human review receives the review-eligible maximisers, unresolved inputs, candidate-specific consequence diagnostics, gate outcomes and evidence/policy version. Approval binds a specific candidate to that version. Evidence acceptance, expiry, revocation and policy revision advance the version and clear approval. Human review is an architectural requirement; automated-decision duties require separate assessment under applicable data-protection law (European Parliament and Council of the European Union 2016, art. 22; Federal Data Protection and Information Commissioner, n.d.).
Context coverage is the fraction of inputs in the manager-approved schema that satisfy their evidence criteria. The schema’s completeness requires external validation. The record reports coverage, itemised gaps, gate outcomes and approval separately (Kurz 2026).
Each transition records the version, actor role, source references and outcome. Revocation or expiry invalidates dependent premises transitively. Evidence acceptance, policy changes and approval are restricted to their declared roles. The reliance record distinguishes a premise’s origin from the act of relying on it (Kurz 2026). Role checks use trusted caller inputs. The event-time invariant assumes serialized transitions; deployment requires an atomic or version-conditioned issuance commit alongside identity, access, tamper-resistance and jurisdiction-specific storage controls.
The transition contract uses a decision version identifying current evidence and policy, a gate predicate , and a recorded human approval . Equation 26 permits issuance only when the current gates pass and a review-eligible maximiser has human approval bound to the same version:
Algorithm 1 routes unresolved mandatory or consequential inputs to permitted questioning or a hold, and requires reassessment before renewed approval after evidence or policy changes.
Algorithm 1: Versioned recommendation procedure for the restricted client selector
Input: Fixed policy, finite candidate and joint-state sets, evidence, question budget
Validate role, evidence dependencies and current policy version
Filter joint states by accepted valid evidence
If evidence is inconsistent or the joint-state set is empty
Hold for evidence resolution
Else if a mandatory gate fails
Request an unresolved permitted input within budget, or hold
Else
Compute , and internally
If
Hold for human evidence resolution, revised alternatives or scope
Else
Assess declared answer partitions and compute
If
Request a consequential permitted input within budget, or hold
Else
Present , consequences, evidence, gaps and version for review
On evidence or policy change: invalidate approval and reassess
On approval: bind the selected review-eligible maximiser to the current version
On issuance: repeat checks and require current bound approval
Proposition 3 (gate preservation). Under the stated transition contract and trusted role boundary, every issuance event satisfies Equation 26 at its event time.
The initial state carries no approval. Evidence and policy transitions clear approval and advance the version. Questioning and refusal leave issuance dependent on a subsequent successful assessment. The approval transition binds a selected review-eligible maximiser to the assessed version. The issuance transition re-evaluates the gates and checks membership in the current eligible set and approval/version equality. Each transition preserves the condition that issuance is possible only through these checks, establishing the event-time invariant by induction over the transition sequence.
Proposition 4 (robust feasible-set expansion). Fix , all utilities and constraints, the active facets and their reference values. For non-empty joint-state sets , the feasible sets defined by Equation 17 satisfy .
Any candidate feasible for every state in satisfies the same constraints over its subset . Taking a utility minimum over that subset weakly raises each lower bound. Every reference-value inequality satisfied under remains satisfied under , so each originally feasible candidate remains feasible.
Recommendation permission also depends on mandatory information and versioned approval. An update can reveal a new liability, alter a constraint or invalidate a source; those changes require renewed assessment under the updated model. The inclusion concerns ; Pareto membership, maximiser membership and review eligibility can change after a restriction of support.
Proposition 5 (undominated issuance). Under the issuance contract, an issued candidate is undominated among current robust-feasible candidates with respect to all declared state/facet utility entries.
Equation 26 requires membership in , which is a subset of and of . Membership in excludes the stated dominance relation by definition, including when every feasible Nash product is zero.
The guarantee depends on supplied utilities and states, whose adequacy for the mandate requires domain assessment.
Results
The analytical constructions establish how information, weights and reference floors affect selection. The portfolio–workflow simulation then measures those effects on the declared client and manager outcomes. Further results concern client-criterion sensitivity, clarification, evidence changes and execution by deterministic and model-based workflows.
The liability construction isolates how undisclosed client information changes selection and permission to issue advice. Client value in this example is the product of a return utility and a liquidity utility, the rule of Equation 18. The synthetic client is an entrepreneur seeking advice on EUR 4 million of investment proceeds after a business sale. Three portfolio alternatives allocate different amounts to assets available within the next 12 months. The amount of a possible payment obligation within that horizon is the context variable. For this illustration, the service’s suitability assessment requires that amount, and each candidate must provide liquid assets . All other legal, data and risk checks are assumed satisfied. 1 fixes the alternatives, with return and liquidity utility scores chosen solely to illustrate the mechanism and reference utilities .
| Portfolio | Liquid assets | Other assets | Return utility | Liquidity utility | Nash product |
| EUR million | EUR million | ||||
| A | 0.4 | 3.6 | 9 | 2 | 18 |
| B | 1.2 | 2.8 | 6 | 6 | 36 |
| C | 2.0 | 2.0 | 3 | 9 | 27 |
B maximises the product at an established obligation of EUR 0.2 million; C alone is feasible at EUR 1.8 million. With unresolved support million, C covers both amounts, while the mandatory-information gate holds issuance. An obligation above EUR 2 million requires revised alternatives or a continued hold. The unresolved input is identical whether the undisclosed obligation is EUR 0.2 million or EUR 1.8 million, while the fully informed winners are B and C. This instantiates Proposition 1: robust C covers both amounts, but an accepted answer is needed to distinguish their client-optimal selections.
The joint construction adds the manager and the delivery workflow to this example, so that both declared evaluations take effect. Client values normalize the portfolio products in 1 by 36. For an operational interpretation of the constructed workflows, uses general case review and uses template-supported review. We assume reusable templates reduce delivery burden for A and C, while B requires additional exception review under . Both workflows are assumed to deliver the same assessed client outcome for a given portfolio. Manager values are stipulated as , where is a normalized burden index: for A, B and C under , and under . Zero denotes the least burden on this illustrative scale. Empirical calibration of the burden index remains prospective.
The established EUR 0.2 million obligation makes all three portfolios financially feasible. The example stipulates that all six pairs pass the base admissibility conditions; domain review must establish which pairs can actually satisfy a mandate’s client-interest duties. The client and manager floors are the values of reference service , namely . The floors retain , and .
3 shows that and are the two efficient pairs meeting both floors; choosing between them requires a client–manager trade-off.
Note. The six portfolio–workflow pairs use the client values and workflow burdens defined in the text. Shading marks the benchmark region; the connecting segment indicates the trade-off between the two efficient discrete alternatives. Source: exact joint-objective construction.
At , Equation 9 selects with score , compared with for . At , wins with , compared with for . The two tie at , each scoring . At the lower weight the selected pair preserves the client benchmark and improves the manager evaluation; at the higher weight it preserves the manager benchmark and improves the client evaluation. The weight makes an explicit choice along the declared trade-off, while the floors restrict permissible sacrifices.
In the terms of Proposition 2, programme Eq. C selects , programme Eq. W selects , and both thresholds equal . At the concession pair of Equation 10 is and , twice the minimum that Equation 11 requires.
4 shows selection moving from to as the client weight crosses , where the two tie.
Note. The three pairs meet both benchmark floors, with all base admissibility gates satisfied. Endpoint values show limits of the specified open weight interval. Source: exact joint-objective construction.
A capacity constraint excluding makes the winner at , even though the excluded alternative has the higher objective score. Missing required information holds every alternative. A further capacity construction admits only and ; each violates one floor, so the joint problem has no feasible solution and holds. The reference service remains the evaluation benchmark while capacity makes it unavailable. Two further capacity settings isolate each floor’s effect on selection. With unavailable and , an objective-only comparison selects at ; the client floor excludes it, and the constrained rule selects at . With only and available and , the objective-only winner is at ; the manager floor excludes it, and wins at . The objective-only calculations isolate the floors’ effects; the operative rule enforces both floors throughout. The eight configurations use constructed manager service evaluations; observed workflow expenditures enter the separate resource comparison below.
The two-pair capacity case gives a concrete review-screen test. Suppose the client confirms both service options and an independent cost assessment establishes their viability. With and hard-admissible, the client comparator before floors is , while the manager floor leaves . Thus , and . With zero omitted client benefit, Equation 14 returns the policy for revision even if policy review has been recorded. Choosing as the agreed reference instead gives floors , selects that pair and reduces both concessions to zero.
A further construction isolates a delivery restriction. Let three services have client–manager values , and , zero floors and . Their scores are , and , so the third is selected. If a documented delivery requirement excludes the first service, it enters admissibility and the second becomes . The remaining client concession is 8, which requires its own supported benefit assessment. These exact constructions also check that zero concession still requires policy review, that benefit evidence identifies the current comparator and policy, and that empty feasible sets hold. They implement the proposed review screen alongside the separate simulation.
The adversarial variations retain both party evaluations and both benchmark floors. Required-input omission produces a hold; accepted support refinement changes the robust admissible set; weight and reference translations change the selected pair. 2 reports these effects using , and . The support-refinement condition uses a separate case policy: certified liability bounds satisfy its required-information gate, and the exact amount is optional. Except where varied explicitly, the reference is , values are those of 3, and the weight is . All 15 exact outcomes matched their hand-derived witnesses; the six common configurations matched the retained evaluator.
| Variation | Controlled change | Joint outcome |
|---|---|---|
| Required evidence | Established input becomes missing at . | to hold. |
| Liability support | Accepted support narrows from to million EUR at . | to . |
| Weight translation | Authorized is transmitted as . | to . |
| Reference translation | Authorized reference is transmitted as . | to . |
| Client-value constraint | Add a hard requirement . | excluded; selected. |
| Workflow effect | Reduce from to , holding other values fixed. | fails its client floor; selected. |
| Affine representation | Transform to and to , including benchmarks. | At : tie. At : restored. |
The hard client-value requirement excludes before scalarisation even though it has the higher base weighted score. This checks precedence of an authorized constraint; institution-specific legal review supplies its substantive content. The workflow variation makes client value depend on delivery, relaxing the original example’s invariance assumption. Under the affine transformation, changing from to preserves every pairwise score difference up to a common positive factor . Leaving the weight unchanged instead changes the effective policy. Inputs and preferences in these stress cases are authored.
The simulation illustrates the joint objective on mandates with financial content: both evaluations now depend on financial and service quantities. Constructed mandates, scales and weights make their effects explicit in a worked implementation. For each pair and scenario, define payment-adjusted economic wealth as . Here is terminal financial wealth, scheduled household spending, the liability valued at its payment-date exchange rate, and total unpaid bills, including service charges. Fees and transaction costs remain losses in this accounting measure; funding adequacy is enforced separately. Equation 27 gives the declared client and manager evaluations used in the joint objective:
The four complete evaluation mandates admit ten, five, ten and seven pairs, respectively, after applying hard constraints and both reference floors. Each retains at least two undominated client–manager trade-offs. Both floors admit every hard-feasible pair in all eight original and corrected states; their restrictive effects are assessed in the analytical constructions. At client weights , , and , the business-sale, retirement, currency-liability and concentration mandates select growth/standard, balanced/advisory, growth/specialist and growth/advisory services. These names identify portfolio and delivery choices; the instruction treatments use the same agent architecture.
3 gives the price of these weights by comparing each selection with the solution of programme Eq. C over the same feasible set. With client value scaled to and manager value to , Equation 12 becomes . In the two business-sale states the selected pair is also the client’s best. In the other six states the rule selects the advisory or specialist workflow where the client-best pair uses the standard workflow, and in two of the four mandates a manager euro counts for more than a client euro.
| Mandate | Selected and client-best workflow | (EUR) | Manager gain (EUR) | ||
|---|---|---|---|---|---|
| Business sale | 0.95 | 0.53 | Standard; identical | 0 | 0 |
| Retirement drawdown | 0.90 | 1.11 | Advisory; standard | 1,190 | 3,062 |
| Currency liability | 0.85 | 1.76 | Specialist; standard | 2,264 | 10,381 |
| Concentrated holdings | 0.98 | 0.20 | Advisory; standard | 1,178 | 6,656 |
Note. Original and corrected liability states of a mandate give the same differences. is the difference in the declared client criterion of Equation 27, converted to euros; the manager gain is the difference in net income . Clarification charges are zero in both selections. The portfolio is the same within each comparison.
The reference service holds cash at a fee of 0.4%. Every hard-feasible pair meets both of its floors, so equals the hard-feasible set in these mandates and the floors restrict nothing here. Consequently here, and 3 reports the total as well as the weight concession. The six positive concessions identify the additional client-benefit evidence that Equation 14 would require in a substantive mandate review; the simulation implements the declared weighted selection.
5 shows the client–manager trade-off for the business-sale mandate. Lowering the transmitted client weight from to redirects the choice from growth/standard to growth/specialist, while both remain admissible.
Note. Complete heldout business-sale case, December 2021 decision. Points are the nine portfolio–workflow pairs and the reference service; both axes use the declared outcome scales. Client value includes the forecast risk penalty and service value. The two rings identify the choices under the authorized and altered weights. Each point is a discrete alternative. Source: frozen simulation inputs and exact valuation.
Eight evaluation states combine the original and corrected liability for each mandate. Faithful optimisation of the complete specifications preserves all eight original decisions. Setting the omitted liability to zero produces two recommendations that violate the original liquidity requirement; replacing omitted manager terms with zero hourly cost and 100 available supervisory hours produces two violations of supervision capacity. Transmitting a client weight of changes four optimal choices while each selected pair remains admissible under the original constraints. These are deterministic interventions in specification content. The optimiser follows the received objective in every condition; the original mandate supplies the reference for assessing fidelity. In the terms of Equation 15, the omitted liability acts on the disclosure term, the omitted manager terms and the mistranslated weight act on the translation term, and the execution term is zero under the common deterministic tie rule used in these controls. 6 separates these constraint failures from admissible changes in the authorized trade-off.
Note. Four mandate families, each evaluated in its complete and revised state. Exact optimization acts on each transmitted specification; outcomes are scored against the authoritative mandate. Categories partition the eight states in each row. Client omission substitutes zero liability and produces liquidity violations; manager omission substitutes zero hourly cost and 100 supervisory hours and produces capacity violations. Complete and revised states share their base mandate. Source: frozen heldout representation controls.
The instruction comparison keeps the declared decision content fixed across three forms. Each reaches all 32 prescribed terminal decisions: 28 issuances and four justified holds. Every issuance preserves the declared numerical mandate, current simulated approval and the joint optimum; unsupported proposals and violations of the encoded constraints are zero. Regret under the declared objective is zero for all 84 issuances, alongside 12 required holds. The equal-output controls assign identical original values to every tied choice, so tie selection leaves the attribution unchanged in these records. Professional prose and nested-delegation instructions each use 20 questions; the explicit joint-objective treatment uses 26. Its six additional questions incur EUR 450 in nominal client charges and EUR 672 in manager labour cost. 4 separates these constructed clarification costs from metered model expenditure.
| Measure | Professional prose | Joint objective | Nested delegation |
|---|---|---|---|
| Correct terminal decisions | 32/32 | 32/32 | 32/32 |
| Issued recommendations / required holds | 28/4 | 28/4 | 28/4 |
| Unsafe raw proposals | 0 | 0 | 0 |
| Issued numerical mandate or constraint violations | 0 | 0 | 0 |
| Questions | 20 | 26 | 20 |
| Client question charges (EUR) | 1,500 | 1,950 | 1,500 |
| Manager question labour (EUR) | 2,352 | 3,024 | 2,352 |
| Reported model cost (USD) | 6.3146 | 7.3317 | 6.3162 |
Note. Four constructed mandates, eight dependent information conditions per mandate, one repetition. Objectives, evidence access, arithmetic and approval controls are shared. Each question incurs a constructed EUR 75 client charge and half an hour of manager labour at the mandate-specific rate. All costs include holds; model expenditure is metered separately.
The comparison contains 96 model episodes; reconstruction reproduces their decisions, evidence transitions and economic outcomes. A public-input reference policy reaches all 64 development and evaluation conditions. The common runtime supplies authoritative source ownership and enforces supported premises, admissibility and current approval. The three instruction forms illustrate equivalent execution of the shared decision rule within those controls.
The subsequent market-path calculation keeps each selected pair fixed and replaces the declared exchange-rate and yield paths with the 2022 observations, retaining the three constructed growth scenarios and their probabilities. The client outcome is plus declared service value under the accounting definition above. Each reference service bears the same clarification expenditure as its corresponding episode. Averaging the eight distinct original and corrected mandate states equally gives a client measure EUR 77,967 below the reference service and manager net income EUR 8,816 above it in every treatment. Manager net income follows from fee income less personnel and fixed costs:
7 compares forecast and subsequent-path client outcomes using the same euro measure, relative to the same reference service. The five-year government yield rises from approximately at the decision date to at the end of 2022, exceeding the stipulated adverse-scenario increase of 1.5 percentage points. The reference service holds cash, whose accrual proxy also rises. This identifies a consequential mismatch between the declared market scenarios and subsequent observations. The outcome calculation preserves the constructed growth assumptions, so interpretation concerns that specified combination of historical and authored paths.
Note. One complete heldout case per family; all three instruction treatments select the same pair. Both bars measure payment-adjusted economic wealth plus constructed service value, relative to the reference, in EUR thousands. The forecast bar averages the three declared valuation scenarios. The 2022 bar uses historical FX and yield observations crossed with the same three constructed growth paths and weights (0.25, 0.50, 0.25). The plotted metric excludes the risk penalty used in decision-time client value. Source: frozen calculations and retained counterfactual cash flows.
The representation checks in 5 hold the candidate set fixed while changing the aggregation specification. Duplicating the liquidity facet changes the complete-context products from to and selects C. Setting reference values to leaves B and C with products 12 and 15. Positive affine utility transformations with correspondingly transformed references preserve the winner: each facet surplus is multiplied by a positive constant, so every candidate product receives the same positive multiplier. For an exact duplicate facet, splitting an existing exponent between its copies preserves that facet’s product contribution; the exponent policy remains part of the declared representation.
| Specification | A | B | C | Maximiser |
|---|---|---|---|---|
| Original product | 18 | 36 | 27 | B |
| Liquidity represented twice | 36 | 216 | 243 | C |
| Reference values | Infeasible | 12 | 15 | C |
| Equal-weight utility sum | 11 | 12 | 12 | B, C |
Taking the original two-facet specification as the declared representation of the mandate, the duplication in 5 changes its implemented weighting while retaining every original utility value and financial constraint. The resulting C choice satisfies the duplicated policy; the original policy selects B. The construction separates execution of a specified aggregation rule from preservation of the declared mandate representation. Which representation reflects a particular client’s preferences requires evidence beyond the stipulated utilities.
All 18 constructed aggregation cases retained two to four robust-feasible candidates. Of the 17 equal-weight comparisons, seven had identical winner sets, six had overlapping but unequal sets and four had disjoint sets. Fifteen cases had conflicting rankings among robust-feasible alternatives; these accounted for five identical sets, six overlapping sets and all four disjoint sets. The separately weighted sensitivity also produced disjoint winners. These counts describe related constructions selected to examine policy behaviour.
A strict disagreement occurs with normalized utilities A=, B= and C=, zero references and common feasibility. Nash products of , and select B. Equal-weight sums of , and select A. The shared constraints admit both selections; their relative desirability depends on the client criterion used to assess the chosen scalarisation. Exact witnesses and reference checks reproduced all 18 Nash calculations.
A further construction isolates the uncertainty objective. With fixed zero reference values, candidate X has utility vectors and in two joint states; candidate Y has in both. Both candidates satisfy all constraints. The product of componentwise minima scores X at 1 and Y at 4, selecting Y. The worst joint-state product, , scores X at 9 and Y at 4, selecting X. The two policies protect different quantities even though the joint states are identical. If X incurs zero monetary loss and Y incurs EUR 100 in each state, the unresolved robust choice has EUR 100 regret. Comparing only the singleton winners would miss this clarification opportunity because both states select X. Equation 24 also compares each resolved choice with the unresolved choice, detecting the difference; the executed case asks at a EUR 10 threshold.
In the constructed comparisons, statewise Pareto filtering removed a dominated candidate whose equal robust score had blocked an eligible alternative, and excluded a dominated choice when all Nash products were zero. Assessing answer cells detected a partial answer that changed the selected portfolio although the unresolved and all singleton choices agreed. Candidate-specific eligibility also allowed an unblocked maximiser to proceed while a tied peer awaited clarification. The full numerical witnesses and comparator traces are reported in the supplementary analytical results.
The partial-answer case shows why the answer partition matters. A declared input separates two of three states from the third. In that two-state cell, changing from A to B increases the client criterion by and reduces monetary loss by EUR 100. The rule asks at a client threshold of 1 or a monetary threshold of EUR 10. A further answer identifying either remaining state changes the preferred candidate back to A, with a client-criterion gain of . Selection can therefore reverse as evidence arrives under the componentwise robust criterion.
The client-criterion test also detects consequences at equal monetary loss. Candidate A has utilities and in two states; B has in both, with zero references and zero losses. The unresolved products are 1 and 4, while each singleton selects A with product 100. Resolving the state yields a client-criterion gain of , triggering a question at a threshold of at most 96. The monetary comparator leaves B eligible for review.
The diagnostics also left consequential information unresolved in two constructions. With flat Nash products and equal monetary losses, both diagnostics were zero although each answer changed which candidate was undominated. With two optional binary inputs, each individual diagnostic was zero while joint resolution changed the winner from B, with product 4, to A, with product 100. Zero individual consequences therefore leave both dominance changes and joint materiality possible. The supplementary witnesses give the utilities and answer partitions.
When an optional answer would restore an empty robust-feasible set, the procedure held for human evidence resolution. Its diagnostics require a feasible current maximiser, so the autonomous question request was rejected. Externally accepted evidence for either answer restored review of the corresponding candidate.
The finite procedure passed 18 checks comprising 26 constructed traces and 14 issuance events. Evidence withdrawal cleared approval even when an irrelevant source or a second independent corroborating source left mandatory support intact. Withdrawal of the sole required source held the decision until accepted replacement evidence restored support and a new approval authorized issuance. Expiry produced the same reassessment under the explicitly advanced logical clock; earlier issuance events remained in the record. Role checks rejected unauthorized evidence, policy and approval operations at the trusted application programming interface (API) boundary. These results establish conformance to the supplied evidence and feasibility labels.
Applying two Swiss service policies to the same allocation states changed permission to proceed. Knowledge and experience and all other relevant service inputs are stipulated as established. The portfolio-based policy classifies the material liability as required financial information; the transaction-specific policy limits required inputs to knowledge and experience. Both retain robust candidate C with zero unmet-liquidity loss, while the portfolio-based policy asks and then holds after refusal and the transaction-specific policy permits review. The difference follows from the declared applicability mapping and conservative institutional hold, with the financial state set held fixed.
6 contrasts the mandatory amount with an optional communication-format preference that leaves utilities and constraints unchanged.
| Condition | Available context | Model outcome |
|---|---|---|
| Complete | Required information established; million. | B proceeds to human review. |
| Missing pivotal input | Required amount unresolved between 0.2 and 1.8 million. | Ask and hold; reassess after an adequate answer. |
| Missing optional context | million established; communication-format preference absent. | B proceeds to review; optional gap remains recorded. |
| Required disclosure withheld | Client declines to supply the necessary obligation information. | Hold and refer to the advisor; mandatory checks remain active. |
A currency mismatch illustrates how a changed valuation can invalidate an earlier liquidity assessment. Assume liquid euro assets of EUR 1.2 million and an unhedged CHF 1.4 million payment obligation, with interest, taxes and execution costs omitted. Each date values the same assumed exposure separately. If is the observed number of Swiss francs per euro, its euro cost is , and the liquidity shortfall is . 7 reports the calculation using observed ECB reference rates (European Central Bank, n.d.-b).
| Date | CHF per EUR | Liability in EUR | Shortfall in EUR |
|---|---|---|---|
| 14 January 2015 | 1.2010 | 1,165,695.25 | 0.00 |
| 15 January 2015 | 1.0280 | 1,361,867.70 | 161,867.70 |
| 16 January 2015 | 1.0128 | 1,382,306.48 | 182,306.48 |
Using the 14 January valuation on either subsequent date would retain an apparent surplus despite the shortfall at the current rate.
The workflow comparisons test whether supplied instructions produce the specified questions, holds and recommendations from partial evidence, while recording delivery resources. All workflows reached the 36 reference decisions in the structured pilot (8), with consistent selections across repetitions. Each made 18 question attempts. Unsafe proposals, issued violations, unnecessary holds and schema failures were zero.
| Measure | Deterministic | Single agent | Multi-agent |
|---|---|---|---|
| Correct terminal decisions | /36 | /36 | /36 |
| Correct permitted recommendations | |||
| Required holds | |||
| Unsafe raw proposals / proposals | /24 | /24 | /24 |
| Issued violations / issued outputs | /24 | /24 | /24 |
| Reported model cost (USD) |
The multi-agent workflow incurred 3.18 times the single agent’s reported model cost. The deterministic workflow used the shared calculator and incurred zero model charges. Implementation, local computation and professional review costs fall outside this accounting boundary.
The retirement construction sets assets at EUR 300,000 and annual pension income at EUR 24,000. Essential expenditure is EUR 5,000 per month and a separate EUR 60,000 payment falls within a three-year reserve horizon. Equation 29 gives the required liquid reserve :
The second construction separates the investment’s price change from its complete cash flows. Equation 30 calculates available assets and the net investment result :
9 gives the four issuance and three hold conditions in each narrative family. The withdrawal grade requires valid initial proposal and approval before evidence loss; a blanket initial hold fails that sequence.
| Information variant | Required evidence transition | Terminal outcome |
|---|---|---|
| Complete evidence | Current records support calculation and approval. | Issue |
| Two jointly missing inputs | Obtain both required records; either adequate question order is accepted. | Issue |
| Refusal | A required input remains unresolved after the client declines. | Hold |
| Resolvable conflict | Signed reconciliation identifies the superseded assertion. | Issue |
| Unresolved conflict | The reconciliation route returns unavailable. | Hold |
| Withdrawal after approval | Evaluator invalidates the source dependencies and approval; workflow responds to the changed state. | Hold |
| Missing presentation preference | Financial inputs remain complete and unchanged. | Issue |
Under fixed responses, all three interfaces reached the required outcomes in all 14 cases; reconstruction reproduced those outcomes.
The initial narrative comparison yielded 42, 30 and 22 correct decisions for deterministic, single-agent and multi-agent workflows, respectively (10). Model-workflow rates were 71.4% and 52.4%; correct recommendations numbered 14 and seven out of 24 issuance cases. Each workflow asked 30 relevant questions and zero redundant questions.
| Measure | Deterministic | Single agent | Multi-agent |
|---|---|---|---|
| Correct terminal decisions | 42/42 | 30/42 | 22/42 |
| Correct permitted recommendations | 24 | 14 | 7 |
| Correct required holds | 18 | 16 | 15 |
| Failed workflows | 0 | 12 | 20 |
| Inadmissible raw proposals / proposals | 0/54 | 1/37 | 0/20 |
| Reported model cost (USD) | 0.0000 | 7.7222 | 13.3133 |
The calculator accepted exactly the required extracted fields and read remaining quantities from the public fixed inputs. Additional fields caused all rejected calculator calls, while repeated successful calculations also consumed the budget. Eight single-agent workflows and 16 multi-agent workflows exhausted the 24-call calculator budget. Response-format errors caused three further single-agent failures and four multi-agent failures. The remaining single-agent failure involved a missing matching calculation. These results measure calculation-interface adherence alongside document interpretation and evidence control.
One single-agent issuance attempt selected the feasible, currently approved R90 alternative but attached an unsupported source link to the fixed sale proceeds. Its premise set also lacked a matching successful calculation. The adapter blocked the attempt; scoring retained it as an inadmissible raw proposal. The enforced issuance invariant recorded zero violations for all workflows. The deterministic, single-agent and multi-agent workflows completed six, four and three of the six withdrawal trajectories, respectively. The other five model attempts failed before initial approval.
Six response-format failures lacked complete final usage totals, two in the single-agent workflow and four in the multi-agent workflow, affecting five case repetitions. Reported costs cover 37 matched records per workflow, while decision analysis includes all 42. The multi-agent cost was 1.72 times the single-agent cost within that subset (10), including metered failures and required holds.
The revised narrative interface completed all 126 assignments, with 42 correct decisions and all six withdrawal trajectories per workflow (11). This revision changed extraction and calculator interaction under the retained case policies; the Pareto and answer-partition rules in Algorithm 1 were assessed separately in deterministic constructions. Failures, validation errors and enforced gate violations were zero. All model proposals were admissible at their evidence versions. Each workflow asked 30 relevant questions and zero redundant questions. Every calculator event returned a single feasible candidate, so these results concern extraction and evidence control.
| Measure | Deterministic | Single agent | Multi-agent |
|---|---|---|---|
| Technical completion | 42/42 | 42/42 | 42/42 |
| Correct terminal decisions | 42/42 | 42/42 | 42/42 |
| Correct permitted recommendations | 24 | 24 | 24 |
| Correct required holds | 18 | 18 | 18 |
| Failed workflows | 0 | 0 | 0 |
| Inadmissible raw proposals / proposals | 0/54 | 0/54 | 0/54 |
| Reported model cost (USD) | 0.0000 | 2.1314 | 5.6348 |
The revised comparison has complete usage accounting for all 42 records per workflow. The multi-agent workflow incurred 2.64 times the single agent’s reported model cost, including the costs of required holds (11). Both used the shared arithmetic-receipt interface. The live run required zero correction attempts.
Independent reconstruction checked all 252 narrative records and their event chronology. The initial run contained 24 deterministic, 14 single-agent and seven multi-agent issued outputs; the revised run contained 24 per workflow. Across these 117 outputs, the checker found zero violations of the supplied financial, evidence, provenance and approval requirements. These results assess the retained emissions under the authored requirements; domain validity remains with the case-policy review.
8 contrasts the initial conformance gaps with the revised 42/42 reference outcomes for every workflow. Revised model expenditures across all 42 assignments are USD 0, USD 2.13 and USD 5.63 for deterministic, single-agent and multi-agent execution. On these two observed dimensions, the deterministic workflow strictly dominates both model arrangements and the single agent strictly dominates the multi-agent arrangement: conformance is equal and model expenditure is lower. This descriptive comparison includes correct recommendations and required holds and retains the experimental grammar, supplied policy and resource-accounting boundary.
Note. Panel (a) shows observed conformance in the initial and revised narrative comparisons; panel (b) shows revised model expenditure with complete accounting. Sources: retained narrative comparisons summarised in Table 10 and Table 11.
Discussion
The simulations draw attention to choices made before an agent begins to optimise. In the constructed cases, an omitted liability led to a liquidity shortfall, omitted manager terms breached supervision capacity, and a mistranslated weight changed the selected service. Duplicating a client facet also changed the objective’s behaviour. These examples suggest a practical reason to review disclosure and mandate translation alongside execution. Proposition 1 gives the corresponding information limit: selecting a client-optimal outcome across states with disjoint optima requires complete decision inputs that distinguish those states.
Nested delegation offers a way to organise that review around the manager’s dual role. The manager agrees the mandate with the client and translates it into instructions for the system. In the proposed framework, mandate review addresses the objectives, scales, service terms and evidence criteria; execution review checks their application and current authorization. Making those responsibilities explicit could help reviewers trace a questionable recommendation to the choices or information on which it depends.
In the joint objective, benchmark floors preserve both declared evaluations relative to a reference service, and positive weights select a Pareto-efficient pair among the alternatives satisfying the constraints and floors. Law and mandate determine those alternatives before scalarisation. This ordering gives a place in the specification to MiFID II Article 24’s client-interest duties and Article 54(1)’s retained responsibility for automated suitability assessment (European Securities and Markets Authority, n.d., para. 1 and 10; European Commission 2017, art. 54(1)). The choice of utility scales, benchmarks and commercial criteria still calls for an assessment against those duties. Even the interpretation of depends on the scales: rescaling manager value while holding the weight fixed changes the effective trade-off.
The manager-floor example makes the stakes of choosing a reference service concrete. Preserving manager value at selects with client value , while excluding with client value and manager value . The weight concession is zero; the benchmark and total concessions are . Under the proposed review rule in Equation 14, that concession calls for policy review and supporting client-benefit evidence. A necessary delivery restriction instead enters admissibility and changes the comparator. Client-confirmed preferences, agreed service terms and domain assessment provide grounds for accepting the policy or revising the reference service.
The economic simulation illustrates a similar tension using constructed fees, personnel costs, supervision hours and service-quality values. In six of eight states, the declared weight selected a higher service tier than the client-best pair, including with a client weight close to one. The floors were nonbinding in those cases, so the reported client amounts equal the total concessions. In the 2022 simulation, expected benchmark preservation also coexisted with a lower subsequent client outcome and higher manager net income. That euro comparison applies the same household-payment adjustment and service valuation to both services; it measures a different quantity from the risk-adjusted decision criterion. The example suggests that review should extend to the forecasts and valuations supporting the mandate.
Service classification matters as well. In the Swiss illustration, the common hold policy adds an institutional restriction to the mapped FinSA information duty; the EU necessary-information gate follows Article 54(8) (Swiss Confederation 2018, arts. 11–14; European Commission 2017, art. 54(8)). A candidate can remain financially feasible while a change in the required-information classification changes permission to proceed. Recording the basis of each restriction could make such decisions easier to explain. The same applies to the distinction between evidence sufficiency and permission to collect it: the specification records both the required information and its permitted acquisition routes (European Securities and Markets Authority 2016; European Parliament and Council of the European Union 2016, arts. 5–6; Swiss Confederation 2020, arts. 6–7).
Clarification raises further questions about how the mandate handles uncertainty. Statewise Pareto filtering removes the demonstrated dominated alternatives, including the zero-product case. Candidate-specific assessment allows an eligible maximiser to proceed while another awaits clarification. The answer-partition test detects the supplied partial-answer consequence, and the client-criterion diagnostic detects a gain from changing candidates even at equal euro losses. These properties hold for the declared states, facets, answer partitions and thresholds. The partial-answer example is particularly relevant to mandate design: the componentwise robust comparison can reverse as evidence arrives, even when every fully resolved state initially selects the same candidate. Whether that behaviour suits the client depends on the chosen treatment of uncertainty.
There are several reasons to seek further information. Required gaps govern permission to proceed; optional answers may improve selection, reveal dominance or restore feasibility. The restricted selector evaluates declared conditional gains under a fixed policy. An answer that would restore an empty feasible set still needs human resolution, while a discretionary question about the full joint objective would need assessment against both parties’ evaluations and the workflow. The simulation assigns costs and source owners to questions. Choosing a sequence of questions would require a further account of their costs and possible answers, with probabilities for an expected-value treatment. Calibration on representative mandates could examine consequential omissions, question burden, unnecessary holds and review effort before evaluating a chosen policy on held-out cases.
After a recommendation, evidence changes may call for a fresh assessment. The withdrawal and replacement checks distinguish lost support from surviving independent evidence; the currency example shows how a changed valuation can leave an unchanged obligation underfunded. Such checks could contribute to the testing, monitoring and allocation of responsibility addressed by FINMA (Swiss Financial Market Supervisory Authority FINMA 2024, sec. 2.1 and 2.4). They also give engineers concrete events that should trigger reassessment or renewed approval.
The narrative runs illustrate how implementation choices can affect execution. Rejected inputs and repeated calculations exhausted budgets in the initial interface, and the adapter blocked a proposal carrying unsupported attribution and missing calculation support. Typed extraction, clearer instructions and shared arithmetic receipts subsequently reached all reference outcomes on the same development cases. This experience suggests that the interface deserves attention alongside the decision rule. Because those changes were introduced together and the correction mechanism remained unused, their individual contributions remain open.
In the instruction comparison, professional prose, the explicit joint objective and nested-delegation instructions reached the same prescribed decisions; professional prose also matched nested delegation on question count. All three received the declared trade-off, both floors, source identities and common enforcement. Their agreement illustrates several ways of communicating the same specification. The representation interventions changed its content and produced departures from the original mandate. Recovering missing inputs used a supplied field schema and authoritative source routes. Recognising an omitted field, choosing an appropriate scale or justifying a reference service would call on professional judgment at an earlier stage.
Workflow design raises a related question about the value of additional roles. In the revised narrative comparison, deterministic execution and the single agent achieved the same conformance as the multi-agent arrangement at lower measured model cost. The two facet-labelled calls and coordinator each received the complete policy and common numerical support, and every revised calculator event had one feasible candidate. The observed agreement therefore concerns extraction and evidence-aware sequencing; the analytical constructions examine selection among competing feasible alternatives. Giving specialists complementary information, with budget-matched repeated single-agent checking as a comparison, could help assess when decomposition adds value.
The service workflow and the agent architecture would both need consideration in an applied system. The simulated service tiers differ in fees, quality and supervisory hours, while the instruction treatments hold architecture fixed. The narrative comparison varies architecture and records model expenditure. Relating these choices would require measurements of professional effort, infrastructure cost and delivered service quality. Additional roles may be worthwhile when their information or expertise improves the service enough to justify those costs. The proposed outcome functions offer a place to express and examine that judgment.
Conclusion
This paper formulates the task of a delegated AI system as joint maximisation under binding constraints within a nested client–manager–AI relationship. The formulation represents client and manager interests through separate outcome functions, with explicit constraints, reference floors and trade-off weights. It connects responsibility for the mandate to its translation into an executable decision specification. Concession accounting makes the client consequences of the floors and weights available for review, while the distinction between disclosure, translation and execution helps locate possible departures from the intended mandate.
The analytical constructions and simulation illustrate how this framework can guide calculations, clarification questions and approval controls. Within the simulated settings, changes in represented information, constraints and policy choices affected the available or selected decisions. These observations suggest that making such choices explicit could support more structured mandate design and review. The underlying proposal is independent of a particular model or agent architecture and may provide a useful basis for comparable implementations. Professional specification studies and field evaluations could assess whether, and under which conditions, this approach improves decision quality, oversight and client outcomes.
Limitations and Future Research
The simulation instantiates the joint objective with constructed mandates, preferences, service terms and growth scenarios. Public exchange-rate and yield observations anchor selected market inputs; they supply current retrieval vintages of historical observations. Client evaluation scales, reference service, quality values and managerial commercial interests require domain validation. The source-linked legal analysis supplies the context for institution-specific mandate and processing review. Strategic disclosure, hidden effort and incentive-compatible implementation require behavioural models and observations of how the parties respond to the specification.
Information degradation depends on the disclosure and specification maps and available supplementary evidence. The indistinguishability result assumes identical complete decision inputs and fixed policy primitives. Finite state enumeration, omitted-state discovery and evidence sufficiency bound the procedure’s applicability. Statewise Pareto filtering and the client-criterion trigger depend on declared cardinal utilities. Flat Nash products can conceal an answer-dependent dominance change even at zero clarification thresholds; undominated issuance concerns the current support. Optional questions that restore an empty feasible set in the restricted client selector require human evidence resolution. The clarification policy examines one-input partitions and possible gains. Joint inputs can conceal a selection gain even when every individual consequence is zero; question probabilities, costs and compound partitions remain open. Applying the legal mapping to an actual mandate requires current provider authorization or exemption, client category, territorial facts and processing arrangements, alongside production identity and storage controls.
The structured pilot spans six families sharing three motivating source clusters; the narratives span two source-linked families with controlled grammar. Repetitions measure within-case consistency. All workflows receive authored policies and numerical support. The revised narrative interface was developed from failures on the same cases used for its evaluation. All 114 revised calculator events returned one feasible candidate, comprising 57 R90 and 57 R180 results. The deterministic parser is tailored to the document grammar. Initial model costs cover 37 matched records per workflow because six failures lack final usage totals; revised costs cover all 42. The deterministic policy checks establish properties of the Pareto and partition procedure, while the historical model comparisons retain their original policy scope.
The joint-objective comparison has four base mandates, eight dependent conditions per mandate and one repetition per instruction treatment. Evaluation mandates were reserved from prompt development and share their four templates and weights with the development mandates; an interrupted preliminary execution exposed a numerical grading defect, corrected by standardising monetary inputs to cents before the reported comparison. The final treatment instructions were retained. Source records present explicit numerical fields, source owners and bounded question routes, and the common approval layer rejects unsupported or inadmissible proposals. The observed parity concerns that shared environment. Broader language variation, independently authored mandates and budget-matched controls would assess transfer; related conditions should remain grouped in allocation and uncertainty estimates.
A prospective reviewer comparison would assess the framework against a competent governance checklist covering the same controls. Reviewers would diagnose matched defects in mandate representation, requirements translation, premise support and current approval. Decision accuracy, responsibility assignment, unnecessary holds and review effort would remain separate outcomes. Practitioner responses and independent domain assessment of case labels would establish whether organizing the controls through nested delegation improves professional specification and review. Field evaluation would examine client outcomes after oversight and implementation costs.
A prospective screening extension would specify disclosure menus and test truth-telling and participation conditions using client utilities, disclosure costs, outside options and outcomes of possible deviations (Rothschild and Stiglitz 1976; Mirrlees 1976). Controlled behavioural studies would examine generative moral hazard, the hypothesis that omitted checks can remain concealed behind a plausible recommendation, and its dependence on effort and generation incentives (Holmstrom 1979; Kalai et al. 2025).
Declarations
Corresponding author. Walter Kurz, me@walterkurz.com
Author contributions. All authors contributed equally to this work.
Funding. This research received no external funding.
Conflicts of interest. The authors declare no conflicts of interest.
Data and code availability. The authors retain the source code, constructed mandates, market inputs, recorded responses, experimental configurations and verification instructions. Access to these research records is restricted. Market inputs include source URLs and content hashes.
AI tools. GPT-6.1 Astra and Claude Opus 5.5 were used for the simulation.
Editorial dates. Submitted 06/2026; revised 09/2026; published 10/2026
Licence. CC BY 4.0