Skip to main content
Configure a Domain for actions that change later states and receive delayed outcomes. Begin with the starting declaration, run the query–execute–feedback loop, then adjust only the settings your task requires. Sequential learning is Adapt-1’s episode-aware policy path for state-changing actions and delayed outcomes. This advanced guide explains how to configure and operate it through the Domain API. When Adapt-1 forms the supported state-dependent action-value and delayed-credit structure from evidence, the Discovery documentation refers to that capability as Sequential Discovery. Declare which states and actions belong to the task and how outcomes become rewards. Choose learning settings for its horizon, feedback density, drift, risk tolerance, and available evidence.
Sequential learning supports authored Domain structure and compatible Discovery paths. The application declares the legal actions, reward semantics, and episode boundaries; supported input structure can form from evidence. The acquisition schedule is separate again: sequential state can form inside a zero-start run, form in a separate trajectory acquisition phase, or continue from a separately acquired checkpoint.

1. When to use sequential learning

Enable sequential learning when all of the following are true:
  • the system repeatedly chooses among two or more policies or actions;
  • actions influence later observations;
  • action quality depends on the current state;
  • reward may be delayed, sparse, or distributed over an episode;
  • events can be grouped into episodes and ordered by step;
  • every action can eventually be linked to a measured outcome.
Examples include control, trading, game play, operations scheduling, treatment sequencing, multi-step remediation, and adaptive workflows. Ordinary contextual policy learning uses outcomes attributed to individual decisions in their decision-time contexts. The application can supply an immediate outcome or compute a return before submitting feedback. Enable the sequential path when Adapt-1 should learn delayed value from ordered transitions. Structured transition learning predicts subsequent observable values. These relationships can coexist in a Domain, with separate targets and output contracts.
Sequential learning is an episode-aware model path inside feedback_policy. A Domain can therefore report feedback_policy and structured_transition as its active subsystems while sequential learning is enabled. When the sequential candidate trains and passes validation, confirm that the serving model is sequential_q_mlp using the model diagnostics below.

Choose zero-start learning, separate acquisition, or both

The declared run determines the acquisition label: state retained across episodes of a zero-start run was acquired inside that run. Sequential fitting requires the configured minimum_episodes and reserves complete episodes for validation. Confirm the serving model is sequential_q_mlp; policy admission during an individual episode can also update ordinary contextual evidence. A separate trajectory acquisition phase helps only when the state fields, actions, reward, terminal meaning, and useful action-value relationship transfer into the new run. Use a zero-start task-local scope when each independently redrawn task has a different hidden mapping and the current context does not identify that difference. Historical trajectories use the same event and feedback contract as live trajectories. Replay approved records into fresh Domain state, preserve the true episode and step order, wait for asynchronous fitting, and inspect the model report before freezing. See Choose a learning setup for the schedule and scope decision.

2. The learning contract

Each sequential transition must provide this logical tuple:
The online loop calls:
  1. POST /domains/{domain_id}/query supplies the current context and returns a selected policy plus a decision_id.
  2. POST /domains/{domain_id}/feedback returns the observed transition and reward.
A decision_id binds feedback to the sealed earlier query. Current runtime context sources are a valid decision_id (decision), a valid target_memory_id (memory), or explicit structured context (request). Sequential readiness is downstream of this policy-admission step. The default field mapping is: All five values must be present and correctly typed for a record to be eligible for sequential training:
  • episode_id: string or other stable identifier;
  • step: finite number, normally a zero-based integer;
  • next_state: JSON object;
  • step_reward: finite number;
  • terminal: JSON Boolean (true or false).

Scope rule

Keep the same authenticated owner and domain_id when learning should carry across episodes. Use metadata.episode_id to separate environment episodes. On the hosted API, the bearer token determines the effective session; changing the body session_id does not create a new tenant or learner. Use a new Domain for an independent task-state history, or a separate credential for a separate identity. The examples use session_id: ignored for compatibility. The following declaration uses example parameters for a delayed-reward policy task. Replace the state fields, policies, policy features, and reward limits with values from the application. Numeric settings in the parameter tables are examples; inspect the resolved Domain configuration for the values in use.
Create it with:

Configure the learner

Each selectable hypothesis needs:
  • a name for humans and traces;
  • the same relation for policies competing in one decision;
  • a unique, non-null policy identifier;
  • optional policy_features describing the action itself;
  • optional when, predicts, and falsified_by clauses for interpretable conditions.
policy identifies an executable action. Its policy_features describe attributes available before execution so the learner can generalize between related actions.Good policy features:
Poor policy features:
Action selection must resolve to a declared policy. A policy-less induced hypothesis supplies predictive evidence. Keep correct-action labels outside autonomous structure targets used for selection.
learning.context.feature_paths determines which state fields are visible to contextual and sequential model training and inference.Include fields that are:
  • available before the decision;
  • causally or predictively relevant to action value;
  • represented consistently at query and feedback time;
  • stable in name, unit, and type.
Exclude fields that are:
  • generated after the action;
  • direct encodings of the answer or reward;
  • identifiers with no transferable meaning;
  • timestamps when elapsed time or phase would be more useful;
  • high-cardinality provenance that encourages memorization.
An empty feature_paths list admits eligible structured leaf fields. Review them for identifiers and logging fields, and exclude outcome leakage. Use explicit paths when the approved inputs must remain fixed.context.event_types filters stored events when a query asks Adapt-1 to infer the latest context. Supply context explicitly when the application has the current observation.context.max_samples bounds retained policy examples. For sequential tasks, size it to retain several representative episodes:
For an example budget of 200-step episodes across 10 retained episodes, 2048 or 4096 provides capacity for the recorded transitions. Replay eviction balances episode groups and preserves failure evidence within the retained sample budget.
Policy learning only updates from measured reward. A successful HTTP response does not imply that the feedback changed the learner.Declarative reward componentsEach component supports:Aggregation modes:Example with performance and safety:
Without components, Adapt-1 recognizes numeric values.reward, values.score, or values.utility; Boolean correct, success, or accepted; and recognized binary outcome labels. Fields such as error_distance have no implicit reward meaning. Either declare them as components or compute a normalized reward externally.Reward scale and neutral rewardKeep rewards in [0,1]. The default neutral point is 0.5:
  • below 0.5 is disadvantageous;
  • 0.5 is neutral or unresolved;
  • above 0.5 is advantageous.
For a naturally signed reward, define normalization explicitly. For example, map profit from [-100,100] with min: -100, max: 100, and goal: maximize.Set component bounds explicitly and inspect the reward values actually admitted. Sequential learning also depends on the accumulated return: repeated extreme rewards can saturate bounded return targets and flatten action-value differences. If estimates cluster at 0 or 1, inspect normalization, neutral reward, discount, and n_step before changing selection gates. A tied query alone does not identify the cause.Keep contextual reward and sequential reward_path meanings explicit. If several recognized reward fields are present, follow the feedback field precedence and avoid conflicting signals.Do not invent a success or failure label for an unknown outcome. Use an outcome value accepted by the current feedback schema and provide the numeric field required by the reward declaration. Do not rely on undocumented outcome labels; the live validator is authoritative for accepted enum values.
Sequential parametersThe learner reserves the last portion of ordered episode IDs for validation. Installation requires validation skill to meet the declared threshold, so retained data must include complete episodes for both fitting and validation.Picking discountUse the effective horizon as a guide:
Approximate examples:Very high discount values can propagate noise and make policies hard to distinguish. Start from the shortest horizon that can still reach the delayed consequence of interest.Picking n_step
  • Dense reward every step: start at 1 to 3.
  • Reward after a short action chain: start at 4 to 8.
  • Sparse terminal reward: start at 8 to 20, with eligibility traces enabled.
  • Highly stochastic outcomes: use a smaller value and more episodes.
n_step controls the observed return window before the model bootstraps from the successor state. For example, n_step: 3 sets a three-reward window before bootstrapping; the successor value can include later consequences. Planning depth is configured separately in a transition-planning workflow. See Decisions beyond the next move.
credit_assignment.mode controls synchronous reward attribution in the online policy state. Sequential Q learning separately consumes the episode transitions and returns.With eligibility traces enabled, a later non-neutral step_reward can update earlier neutral transitions in the same episode. Propagation stops at a prior non-neutral signal, a terminal boundary, an unchanged transition, or the minimum weight cutoff.Use mode: none when rewards are immediate and accurately assigned. Use eligibility_trace when the effect arrives later than the responsible action.Do not use a large minimum_weight for long sparse-reward episodes. With discount: 0.95, a cutoff of 0.05 reaches roughly 58 steps; a cutoff of 0.01 reaches roughly 90 steps.
Exploration modesExploration requires allow_exploration: true in the query and uses the mode configured by the Domain. A query can temporarily override the mode with selection_mode.Record the returned mode when inspecting auto; its choice can remain unchanged across successive decisions. Posterior sampling can be reproducible for a sealed decision. Compare the context, admitted evidence, and returned diagnostics when investigating repeated selections.exploration_strength scales UCB uncertainty. 0 removes the UCB bonus. Values near 0.5 are cautious, 1.0 is a reasonable start, and values above 1.0 deliberately favor underexplored actions.Selection gatesSet selection gates according to the permitted deployment risk. Use held-out operational data to choose thresholds after collecting initial coverage. The client must always handle status: "abstained"; abstention is a valid safety output.
Policy transferSet both transfer strengths to 0 when actions or relations have unrelated semantics. Increase action_transfer_strength only when policy features have stable physical or operational meaning. Transfer never crosses owner, Domain, or session scope.Decaydecay_half_life_seconds exponentially reduces the influence of older feedback. 0 disables time decay.Use decay for regime changes, market drift, changing users, or equipment aging. Choose a half-life in application time that smooths ordinary noise and allows obsolete evidence to decay.Examples:
  • stable physical process: 0 or several months;
  • weekly operational drift: several days to weeks;
  • intraday regime changes: tens of minutes to hours;
  • controlled reversal test: long enough to retain evidence, short enough for new evidence to dominate.
Latent belief separates recurring or changing context regimes without requiring the client to provide a regime label.Tuning effects:
  • Lower novelty_threshold: creates regimes more readily; useful for distinct modes, risky with noisy features.
  • Higher novelty_threshold: merges more contexts into one regime.
  • Lower temperature: sharper regime assignments.
  • Higher temperature: blends evidence across nearby regimes.
  • Lower change_point_patience: reacts faster and risks false changes.
  • Higher historical_regime_weight: gives evidence from earlier regimes more influence.
Tune latent belief after verifying feature scaling and reward semantics, including episode linkage.
A fresh Domain starts these policy learners without task-trained state. Eligible feedback updates the online policy posterior synchronously and triggers background candidate fitting during operation when the configured training gates are met. Queries use the available policy state while a candidate is fitted and validated.retrain_interval controls how many newly accepted samples are required after a completed fit. Choose an interval for the acceptable compute load and adaptation delay. Example settings:
  • rapid drift, cheap fitting: 8 to 20;
  • stable task, moderate load: 32 to 100;
  • expensive fitting or very large streams: 100 or more.
Training readiness also depends on minimum_episodes. A fit triggered before enough episodes exist may train the ordinary contextual candidate; once enough valid episodes accumulate, the sequential candidate can be trained and validated.

13. Execute the online loop

Step 1: query before acting

Persist these response fields:
  • decision_id;
  • selection.status;
  • selection.selected_policy;
  • the exact action actually executed;
  • the pre-action context.
If the application overrides Adapt-1’s selection, send feedback for the policy that was actually executed. Do not credit an action that was only proposed.

Step 2: observe the transition

After executing the selected policy, collect:
  • the successor state;
  • the reward components;
  • whether the episode ended;
  • the episode ID and step;
  • any measured delay.

Step 3: write decision-linked feedback

Repeat query, action, and feedback until the terminal transition. The terminal row still needs a valid next_state object and must send "terminal": true. decision_id binds the query → action → outcome loop to sealed pre-action context. Current runtime policy admission can also recover context from target_memory_id or from explicit structured context. In every case, send the executed relation and policy explicitly and verify that credit_assignment.contextual_learning_applied is true. Current admission checks require the relation field to be present and do not validate its value against the bound decision. Send the real relation from the Domain. A successful feedback response without relation can store the record while leaving feedback_policy.sample_count unchanged.

14. Batch feedback without losing sequence semantics

POST /domains/{domain_id}/batch accepts ordered event and feedback operations. Every feedback operation retains its own decision_id, metadata.episode_id, metadata.step, successor state, reward, and terminal flag. Batch feedback when the corresponding decisions and outcomes have already occurred. Preserve each transition identity and order. Continue querying between actions when a new action depends on the previous result. Minimal shape:

15. Read the diagnostics

Request a normal Domain query and inspect: The top-level learning_state.sample_count is an aggregate across active learning subsystems. Use the feedback-policy subsystem count when diagnosing sequential policy learning. Diagnose sequential fitting through learning_state.subsystems.feedback_policy. Its admitted feedback samples supply the episode and training gates for sequential_q_mlp. Check this admission before investigating training thresholds.
Within learning_state.subsystems.feedback_policy.model, report describes a fit. After a candidate is rejected, it can describe that candidate while current_model_type and current_validation_skill identify the retained serving model. A response with status: retained and installed: false can therefore still have an active model. After a successful installation, the report describes the installed fit. trained records a completed fit; inspect installation and retention fields together.The model identifiers here describe fitted estimators available to the configured feedback_policy inference path. Contextual candidates can report extra_trees or mlp_v2; sequential_q_mlp identifies the sequential return path. Candidate validation contributes to selecting the serving estimator. Query model_weight reports that estimator’s contribution for the current context. The Mechanisms page describes how model fitting operates alongside learned task structure, retained evidence, and regime state. report.sample_count describes the fit’s training data and can lag the current learner count.The reason fields describe different decisions:
  • model.reason: why fitting retained the incumbent or rejected a candidate, such as candidate_failed_validation or candidate_objective_mismatch;
  • model.report.selection_reason: model selection, such as robust_model_within_validation_tolerance;
  • query selection.reason: action selection or abstention.
In contextual-policy diagnostics, model_distribution.status: latent_fallback with model_support_unavailable in model_distribution.reasons[] means trained-model support was unavailable for that estimate. Check model_weight and OOD support alongside it. These diagnostics distinguish a completed fit from useful model contribution on the current query.
Verify sequential learning through these diagnostics:
  1. accepted feedback updates the feedback-policy subsystem; inspect its sample count and learner version, allowing for the retained-sample limit;
  2. multiple valid episode IDs are retained;
  3. fitting completes and a validated model is installed or retained;
  4. the serving model type is sequential_q_mlp;
  5. the serving model’s validation skill meets the configured threshold;
  6. sequential expected rewards differ across policies or contexts and contribute to selection;
  7. evaluate frozen exploit decisions on held-out episodes to measure the effect on task performance.

16. Freeze evaluation correctly

To evaluate a frozen learned state:
  1. train only on the training partition;
  2. wait for asynchronous training to complete;
  3. stop all /feedback writes;
  4. query held-out episodes with allow_exploration: false;
  5. use selection_mode: "exploit";
  6. keep the same authenticated owner and domain_id so the trained state remains available;
  7. use unseen episode IDs, and preferably unseen scenarios or entities;
  8. record abstentions in the Adapt-1 score and score application fallbacks separately.
Frozen query example:
Report at least return, policy accuracy if ground truth exists, abstention coverage, selective performance, episode variance, and learning curves by episode. Use multiple seeds or independent streams for stochastic environments.

Advanced configuration

These profiles provide example parameters for evaluation on the application task.Sparse delayed reward
Use when most steps are neutral and terminal outcomes carry the useful signal.Dense short-horizon control
Nonstationary or reversal-prone task
Set the decay half-life in real application time. 86400 is one day and is only an example.High-stakes conservative deployment
Collect exploration data in a safe environment before enabling these deployment gates.Many structurally similar actionsUse informative policy_features, set action_transfer_strength around 0.1 to 0.25, and increase conservative_weight when many actions remain unsupported. Compare against action_transfer_strength: 0 to measure the effect on held-out actions and check for interference between distinct actions.
learning.sequential.bound_transition is an optional non-neural projection for tasks with explicitly bound state objects and action roles. It can estimate how a declared transition action changes a bounded objective and blend that estimate into policy scoring. Most sequential tasks should leave it disabled.It is appropriate when:
  • one context contains multiple indexed entities or objects;
  • a policy selects which entity is the current state and optionally which is the action source;
  • action roles can be declared from policy features;
  • the objective has known numeric bounds;
  • state transformations can transfer between structurally similar bindings.
Enable bound-transition projection when the application can declare selectors, state features, action roles, and objective semantics independently of evaluation labels.

19. Common failure modes

sample_count stays at zero

First distinguish storage from learner admission. A successful feedback response can store the record while leaving feedback_policy.sample_count unchanged. Check:
  • learning.enabled is true;
  • feedback includes a real measured reward;
  • relation is present in the feedback request;
  • the executed policy is present and is one of the declared hypotheses;
  • decision-time context is recoverable from a valid decision_id, a valid target_memory_id, or explicit structured context.
Then inspect credit_assignment.contextual_learning_applied. If it is false, the record did not enter contextual policy learning even if the HTTP request succeeded. Current responses report the context route under credit_assignment.context_source as decision, memory, or request. An unrecognized outcome label does not update policy learning unless a numeric, Boolean, or declarative reward is also present.

Feedback accumulates without an installed sequential model

Check every row for all five configured sequential paths and types. Then check:
  • at least minimum_episodes distinct episode IDs exist;
  • replay capacity retains enough rows from those episodes;
  • training has reached min_samples;
  • Torch or the configured training service is available;
  • asynchronous fitting has completed;
  • validation skill meets minimum_validation_skill.

Every decision abstains at the start

Use allow_exploration: true during data collection. With no policy evidence and exploration disabled, abstention is expected. Keep strict confidence, dominance, and OOD gates disabled until there is enough coverage.

Decisions return selected_policy: null

Only declared hypotheses with non-null policies are executable. Predictive or induced hypotheses without a policy cannot be selected as actions. Verify that all competing policy hypotheses share the requested relation and have unique policies.

Learning becomes worse across episodes

Check for:
  • reward direction or normalization errors;
  • incorrect decision-to-feedback linkage;
  • outcome leakage in context features;
  • reused episode IDs or non-monotonic steps;
  • too-small replay capacity;
  • excessive transfer between unrelated actions;
  • too-high discount or n_step for noisy rewards;
  • validation episodes drawn from a different regime due to lexicographic episode naming;
  • exploration still enabled during evaluation;
  • client fallbacks being scored as if Adapt-1 selected them.

Policy ordering stays flat after training

Inspect validation_skill, sequential_expected_reward, model weight, and OOD status. Common causes are constant rewards, action-independent features, identical policy features, low action coverage, a feature path mismatch between current and successor state, or OOD downweighting.
Use this sequence:
  1. Define the decision. Write one sentence: “Given state X, choose one policy from Y to maximize measured outcome Z over horizon H.”
  2. Declare policies. Give every executable action a unique policy and observable policy features.
  3. Declare pre-action state. Select only fields known before acting.
  4. Define reward mathematically. Normalize every objective to [0,1]; specify tradeoffs and hard constraints.
  5. Define episode boundaries. Establish stable episode IDs, ordered steps, successor state, and terminal semantics.
  6. Estimate the horizon. Pick the shortest discount and n_step that reach the relevant consequences.
  7. Size replay. Retain enough transitions from enough complete and diverse episodes.
  8. Choose exploration. Explore in a safe environment; exploit or abstain in high-risk deployment.
  9. Set validation gates. Require positive held-out skill and keep conservative penalties when action coverage is incomplete.
  10. Instrument the loop. Persist decisions and verify that every executed action receives correctly linked feedback.
  11. Run ablations. Compare sequential enabled/disabled, transfer enabled/disabled, and latent belief enabled/disabled.
  12. Freeze and evaluate. Use unseen episodes with no feedback writes and no exploration.
  • The authenticated owner and Domain stay stable across training episodes.
  • Every policy is unique, executable, and attached to the queried relation.
  • Policy features describe action attributes available before execution.
  • Context features contain no post-action or answer information.
  • Reward is measurable, normalized, and directionally correct.
  • Every feedback row has episode, step, next state, step reward, and Boolean terminal.
  • Every executed action has one valid decision-time context binding: sealed decision_id, valid target_memory_id, or explicit structured context; use decision_id for sealed query attribution when available.
  • Replay capacity covers several representative episodes.
  • Exploration is enabled only where exploration is acceptable.
  • The client handles abstention explicitly.
  • A sequential_q_mlp is installed only after positive held-out validation skill.
  • OOD and confidence behavior is tested before high-stakes deployment.
  • Evaluation is frozen, held out, and free of feedback writes.
  • Results include variance, coverage, abstention, and learning curves.

Sequential Discovery

Start with the Discovery behavior, public boundary, and minimal contract.

Discovery overview

Review Transition, Structure, and Sequential Discovery together.