USL10 Paper Portfolio
USL-06 Belief Validity · Evidence · Acting AAAI-27 abstract registered · forum linked public contract: drafting

Learn the Unit While Acting: Active Identification and Sequential Decision-Making

How should an agent act when actions affect reward, dynamic state and future evidence about a persistent actual unit?

venue AAAI-27 Main Track topic RU: Decision/Utility Theory & Sequential Decision Making state abstract registered / writing owner gong · CausaClaw × DiscoSeed OpenReview forum ↗

基本研究问题

这篇论文独立可证伪的问题是什么;它与共享 ontology 中其他九篇不重叠的部分。

How should an agent act when actions affect reward, dynamic state and future evidence about a persistent actual unit?

核心 claim
Sequential decisions require a history-conditioned joint belief over the Unit Selection Variable realization and evolving state, with error or regret decomposed into unit-identification and action components rather than treating the learner as the selector.
如果成立会改变什么
Diagnostic actions can be valued for learning which individual was selected, not only for reducing uncertainty about the current state.
最强 reviewer 反对 / kill signal

反对:The unit is merely a hidden MDP state, task ID or bandit arm.

Kill signal:merge into standard POMDP work if `U` has no independently testable persistence or query semantics beyond hidden state.

基本思路

论文如何回答这个问题:形式化对象、论证 spine 与所需证据。

USL-01 type contract

The Unit Selection Variable $U\sim\Pi$ is sampled once for an episode under the task-declared Population law and performs individual selection; on $\{U=u\}$ the realized individual $u$ persists across time. The learner does not select that individual. Before action $A_t$, admissible history $\mathcal H_t$ first conditions the world law $P(U\in du\mid\mathcal H_t)$; the learner forms $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$. If state $S_t$ is hidden, the learner generally needs a joint belief $B_{\phi,t}(du,ds\mid\mathcal H_t)$; posterior independence is not a default. A separately typed joint fixed-individual kernel such as $M_\theta(ds_{t+1},dr_t,do_{t+1}\mid s_t,a_t;u)$ retains randomness at fixed $u$ without silently imposing conditional independence between next state, reward and observation. Factoring this kernel is an application assumption; a complete episode law must also declare its initial-state law and policy. Candidate $A_t$ and current $R_t$ are not pre-action evidence; after execution their factual realizations may enter $\mathcal H_{t+1}$. Actions change state/evidence laws but do not reselect $u$ unless a different protocol is explicitly modeled. USL06 owns sequential control under an evolving state/reward episode law. Pure diagnostic evidence-channel design for a family of future queries belongs to USL08 unless controlled dynamics and cumulative reward are scientifically central.

论证 spine

  1. Episode-level Population selection, persistent realized individual and evolving state.
  2. World conditional $P(U\in du\mid\mathcal H_t)$, learner approximation $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$ and learner joint belief over individual and hidden state.
  3. A separately typed joint stochastic next-state/reward/observation kernel at fixed $u$; any factorization is an application assumption.
  4. Pre-action evidence, action value, information value and evidence cost.
  5. Unit-identification/belief error and action regret.
  6. Diagnostic-action environment and matched POMDP/bandit baselines.
  7. Safe defer, support constraints and offline-policy limits.

所需证据

  • One policy-separation example driven by unit information value.
  • Matched Bayes-adaptive/POMDP and latent-bandit baselines.
  • Ablations collapsing `U` into state or task ID.
  • A joint-belief baseline showing whether any claimed gain survives without assuming posterior independence between $U$ and $S_t$.
  • A complete episode-law declaration including the initial-state law and policy, with a joint fixed-$u$ kernel or an explicit and testable factorization.
  • A timing audit that excludes candidate actions and current rewards from pre-action evidence and verifies that actions do not silently reselect the individual.
  • A separate calibration/estimation audit for $Q_\phi(du\mid\mathcal H_t)$ against $P(U\in du\mid\mathcal H_t)$ rather than treating the learner as the selector.
已注册 abstract(点击展开)

We formulate sequential decision-making in which pre-action history $\mathcal H_t$ first conditions the world law $P(U\in du\mid\mathcal H_t)$ and the learner forms $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$, while maintaining a separate belief over dynamic state. Diagnostic actions are valued for both immediate reward and their information about which unit was selected.

当前进展

状态只记录可验证 delta:claim、SOTA opponent、theorem、experiment、manuscript 或 owner decision。

当前状态

AAAI-27 abstract registered; full paper drafting.

下一写作门

formalize episode-level Population selection, persistent realized $u$, dynamic $S_t$, pre-action history $\mathcal H_t$, world conditional $P(U\in du\mid\mathcal H_t)$, learner approximation $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$, learner joint belief, stochastic transition/reward/observation kernels and action/evidence costs.

下一证据门

show an information-seeking action selected only by the unit-aware objective.

48h 最小实验

construct a small environment with matched current observations and separately typed fixed-$u$ transition/reward kernels; compare myopic, ordinary belief-state and selector-aware policies without granting oracle attribution.

Closest SOTA opponents

Bayes-adaptive MDPs/POMDPs; contextual and latent bandits; active system identification.

canonical 文件

research-questions/USL06-learn-the-unit-while-acting/seed.md
papers/USL06-learn-the-unit-while-acting/paper.md

进展日志

2026-07-29 · AAAI-27 abstract registration 由 direct OpenReview forum link 验证;full paper drafting 启动;本页面建立为持续 review 面。

后续每次可验证 delta 追加在此;canonical 状态以 seed.md / paper.md / submission-ledger 为准,本页由 build_paper_pages.py 重新生成。