How should an agent act when actions affect reward, dynamic state and future evidence about a persistent actual unit?
这篇论文独立可证伪的问题是什么;它与共享 ontology 中其他九篇不重叠的部分。
How should an agent act when actions affect reward, dynamic state and future evidence about a persistent actual unit?
反对:The unit is merely a hidden MDP state, task ID or bandit arm.
Kill signal:merge into standard POMDP work if `U` has no independently testable persistence or query semantics beyond hidden state.
论文如何回答这个问题:形式化对象、论证 spine 与所需证据。
The Unit Selection Variable $U\sim\Pi$ is sampled once for an episode under the task-declared Population law and performs individual selection; on $\{U=u\}$ the realized individual $u$ persists across time. The learner does not select that individual. Before action $A_t$, admissible history $\mathcal H_t$ first conditions the world law $P(U\in du\mid\mathcal H_t)$; the learner forms $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$. If state $S_t$ is hidden, the learner generally needs a joint belief $B_{\phi,t}(du,ds\mid\mathcal H_t)$; posterior independence is not a default. A separately typed joint fixed-individual kernel such as $M_\theta(ds_{t+1},dr_t,do_{t+1}\mid s_t,a_t;u)$ retains randomness at fixed $u$ without silently imposing conditional independence between next state, reward and observation. Factoring this kernel is an application assumption; a complete episode law must also declare its initial-state law and policy. Candidate $A_t$ and current $R_t$ are not pre-action evidence; after execution their factual realizations may enter $\mathcal H_{t+1}$. Actions change state/evidence laws but do not reselect $u$ unless a different protocol is explicitly modeled. USL06 owns sequential control under an evolving state/reward episode law. Pure diagnostic evidence-channel design for a family of future queries belongs to USL08 unless controlled dynamics and cumulative reward are scientifically central.
We formulate sequential decision-making in which pre-action history $\mathcal H_t$ first conditions the world law $P(U\in du\mid\mathcal H_t)$ and the learner forms $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$, while maintaining a separate belief over dynamic state. Diagnostic actions are valued for both immediate reward and their information about which unit was selected.
状态只记录可验证 delta:claim、SOTA opponent、theorem、experiment、manuscript 或 owner decision。
AAAI-27 abstract registered; full paper drafting.
formalize episode-level Population selection, persistent realized $u$, dynamic $S_t$, pre-action history $\mathcal H_t$, world conditional $P(U\in du\mid\mathcal H_t)$, learner approximation $Q_\phi(du\mid\mathcal H_t)\approx P(U\in du\mid\mathcal H_t)$, learner joint belief, stochastic transition/reward/observation kernels and action/evidence costs.
show an information-seeking action selected only by the unit-aware objective.
construct a small environment with matched current observations and separately typed fixed-$u$ transition/reward kernels; compare myopic, ordinary belief-state and selector-aware policies without granting oracle attribution.
Bayes-adaptive MDPs/POMDPs; contextual and latent bandits; active system identification.
research-questions/USL06-learn-the-unit-while-acting/seed.mdpapers/USL06-learn-the-unit-while-acting/paper.md
后续每次可验证 delta 追加在此;canonical 状态以 seed.md / paper.md / submission-ledger 为准,本页由 build_paper_pages.py 重新生成。
阅读后直接在 Discord 写作工作区对应篇目下留言:点出 claim / 思路 / 进展中需要改的具体位置,DiscoSeed 与 CausaClaw 会把决定回填到 canonical 文件并重新生成本页。
进入 Discord 写作工作区 ↗ 返回十篇总览