USL10 Paper Portfolio
USL-09 Generation · Generalization · Personalization AAAI-27 abstract registered · forum linked public contract: drafting

In-Context Unit Selection for Personalized Foundation Models

Can foundation-model context identify a persistent real-world user or entity without making sample representations stochastic?

venue AAAI-27 Main Track topic NLP: Prompt Engineering & In-Context Learning state abstract registered / writing owner gong · CausaClaw × DiscoSeed OpenReview forum ↗

基本研究问题

这篇论文独立可证伪的问题是什么;它与共享 ontology 中其他九篇不重叠的部分。

Can foundation-model context identify a persistent real-world user or entity without making sample representations stochastic?

核心 claim
Context observations can be evidence about a persistent pre-existing unit, and the inferred unit belief can be reused across unseen queries while token/sample representations remain deterministic.
如果成立会改变什么
Personalization would separate who the context is about from the content of the current prompt and from transient conversational state.
最强 reviewer 反对 / kill signal

反对:This is user embedding, retrieval, prefix tuning or latent in-context learning.

Kill signal:downgrade if gains disappear when content leakage is removed or if the belief cannot support a reusable cross-query audit.

基本思路

论文如何回答这个问题:形式化对象、论证 spine 与所需证据。

USL-01 type contract

The Unit Selection Variable $U\sim\Pi$ selects a persistent Individual under the task-declared Population law $\Pi$. An observed context array is an encoding of admissible information $\mathcal O$ and conditions $P(U\in du\mid\mathcal O)$; the context does not perform selection. The model forms a learner-specific $Q_\phi(du\mid\mathcal O)\approx P(U\in du\mid\mathcal O)$ and evaluates that approximation rather than defining the world law. Known and ambiguous referents are different regimes. With reliable attribution, user_id determines the task-declared unit value $u_k$, and $P(U\in du\mid \mathrm{ID}=k,\mathcal O)=\delta_{u_k}(du)$. This is a legitimate known-ID instantiation of the Unit primitive, not merely a computational baseline. A high-dimensional vector or embedding-row lookup is one common instantiation of this boundary, not the definition of $u$ or a requirement on the unit space. Any non-degenerate belief over a separately typed user state or latent property $Z_{u_k}$ is then property/state uncertainty, not which-individual uncertainty. A non-degenerate learner unit belief $Q_\phi$ is licensed only when the referent itself remains ambiguous; it is not the world selector. The identity-bearing $u$ enters the response contract directly. Within the task-declared unit space, it is the formal Individual and selects the whole response-law member rather than merely serving as an extra computational input; this is a modeling convention, not a metaphysical claim. An implementation may encode $u$ computationally, but that encoding neither defines nor replaces the unit. If personalization also uses $Z_u$, type it as an additional property or state rather than a mandatory second embedding. Conditional on $\{U=u\}$, a response law such as $K(dy\mid q,\mathcal O;u)$ may remain non-degenerate: deterministic context or sample representations do not remove event/exogenous output randomness. Repeated sessions require an explicit linkage/observation law, and a current candidate prompt or unobserved target is not automatically part of $\mathcal O$.

论证 spine

  1. Unit Selection Variable $U$, task-declared Population law $\Pi$, formal Individual $u$, optional derived property, content representation and transient state are different objects.
  2. Context as realized evidence: world conditional versus learner inference.
  3. Reliable-ID regime: `user_id` determines the task-declared unit value $u_k$, so $P(U\in du\mid \mathrm{ID}=k,\mathcal O)=\delta_{u_k}(du)$; this is a legitimate known-ID Unit reduction, while a vector or embedding-row lookup is only one common instantiation and uncertainty over optional $Z_{u_k}$ or user state remains property/state uncertainty.
  4. Ambiguous-referent regime: a non-degenerate learner unit belief $Q_\phi(du\mid\mathcal O)$ approximates the corresponding world conditional and is reused without treating the candidate query as evidence.
  5. Repeated-user benchmark with declared linkage and identity/content decoupling.
  6. Comparison with embeddings, retrieval, memory and in-context meta-learning.

所需证据

  • A leakage-resistant repeated-user split.
  • Cross-query/session improvement beyond matched retrieval and user embeddings.
  • A conflict case where a belief distribution is safer than point personalization.
  • Separate calibration of $Q_\phi(du\mid\mathcal O)$ against the declared world conditional or selector truth; predictive success alone is insufficient.
  • A trusted-ID control in which the world which-unit conditional is Dirac, separating selector error from non-degenerate posterior uncertainty over a known user's property or state.
  • A matched known-ID control recognized as the Dirac Unit reduction (`user_id` $\to u_k$), including a common vector-instantiated user-embedding implementation without treating that embedding as the definition of a unit, plus any optional latent response property $Z_u$; the same $u$ must select the whole response-law member, with an explicit repeated-session linkage law.
  • A fixed-$u$ stochastic response/noise audit showing that deterministic context or sample representations do not make the selected individual's outputs deterministic.
已注册 abstract(点击展开)

We treat context observations as evidence about a persistent real-world user or entity that exists independently of the model. Within a task-declared unit space, the formal value $u$ denotes that Individual and selects a whole query-to-response law; this is a modeling convention rather than a metaphysical claim. The model retains deterministic sample representations while forming and reusing a belief over which actual unit is selected.

当前进展

状态只记录可验证 delta:claim、SOTA opponent、theorem、experiment、manuscript 或 owner decision。

当前状态

AAAI-27 abstract registered; full paper drafting.

下一写作门

separate Unit Selection Variable $U$, task-declared Population law $\Pi$, formal Individual $u$, admissible observed context, current conversation state, sample representation, any optional derived property $Z_u$, and output distribution; treat a verified `user_id` $\to u_k$ lookup as a legitimate known-ID Dirac Unit reduction, with embedding-row lookup only one common vector-valued instantiation, state when repeated-session linkage is observed versus inferred, and keep the candidate prompt/current target outside selector evidence unless already factual and admissible.

下一证据门

run a repeated-user split with identity/content decoupling.

48h 最小实验

construct repeated-user episodes with identity-confounded content, held-out query types and conflicting context; compare pooled, embedding, retrieval and belief-based methods.

Closest SOTA opponents

personalized language models and user embeddings; in-context Bayesian/meta-learning; retrieval-augmented memory systems.

canonical 文件

research-questions/USL09-in-context-unit-selection/seed.md
papers/USL09-in-context-unit-selection/paper.md

进展日志

2026-07-29 · AAAI-27 abstract registration 由 direct OpenReview forum link 验证;full paper drafting 启动;本页面建立为持续 review 面。

后续每次可验证 delta 追加在此;canonical 状态以 seed.md / paper.md / submission-ledger 为准,本页由 build_paper_pages.py 重新生成。