USL-01Unit / sample primer
USL-01 deep technical primer · object typing

U = u
到底是什么意思?

Population \(U\) 随机选择“是哪一个 individual”;event \(U=u\) 固定这个 individual。Factual evidence \(\mathcal O^F\) 一次形成 \(Q_\phi(du\mid\mathcal O^F)\),多个 \((x^Q,c^Q)\) queries 复用它;query 不重新进入 Unit Abduction。

public owner-review preview · not a formal releasev1.6 · 2026-07-28 direct-address-response-content contractlocal source v0.15 · public contract v0.12verified main PDF v0.14 · supplement v0.9noindex · navigation refreshed 2026-07-31
The whole entry formula

\(U:\Omega\to\mathcal U,\ U\sim\Pi;\quad \mathcal O^F\to Q_\phi(du\mid\mathcal O^F);\quad \widehat P(dy^Q\mid\mathcal O^F;x^Q,c^Q)=\int P_\theta(dy^Q\mid x^Q,c^Q,u)Q_\phi(du\mid\mathcal O^F)\)

ABDUCTION · \(Q_\phi(du\mid\mathcal O^F)\)Factual evidence conditions the already-realized Unit and forms one learner belief.
RESPONSE · \((x^Q, c^Q, u)\)Alternative queries reuse that belief; direct response information is typed as \(c^Q\), not the whole evidence bundle.
\(\Pi\) is the unit-space base law; query does not reselect or re-abduct the Unit.
01 · type system

先分清 Population、individual、evidence 与 learner

当前 foundational choice 只有一个 task-declared unit primitive:\(u\in\mathcal U\) 表示 persistent referent,并以语义方式索引 \(f(x;u)\) / \(P_\theta(dy\mid x;u)\);\(\mathcal U\) 不必是数值空间。关键不在给任意对象改名,而在同时声明 Population-level selector、fixed realized individual、world conditional、learner belief 与 fixed-unit response。

对象记号 / 类型职责它不是什么
Population selector\(U:\Omega\to\mathcal U,\ U\sim\Pi\)随机回答“从 Population 中选中哪一个 individual”\(\Pi\) 不直接规定 dataset sampling
Fixed individual\(u\in\mathcal U,\ \{U=u\}\)固定/表示 definite persistent referent,以语义方式索引 \(P_\theta(dy\mid x;u)\)不是一条已完成的数值向量,也不是 trainable \(u_\theta(k)\)
Evidence\(\mathcal O\)回答前可用的 admissible observed event / information不是任意输入数组或候选 query
World conditional\(P(U\in du\mid\mathcal O)\)给定 Population information 时,已由 \(U\) 实现的 individual 的总体条件分布不是 learner 自己的输出
Learner belief\(Q_\phi(du\mid\mathcal O)\)learner 对 world conditional 的 epistemic belief,以其为目标或规范对象不是 fixed-u response noise
Learner response model\(P_\theta(dy\mid x^q;u)\)对 fixed-individual population response law 的近似不是 which-unit uncertainty,也不自动等于 world law

Observed-attribution special regime:只有数据已经可靠揭示 row 属于哪个 individual 时,才可以写 \(\mathcal D_U^{\rm obs}=\{(x_i,y_i;u_i)\}_{i=1}^N\)。此时 \(u_i=u_j\) 表示两条 samples 归属于同一个 individual;这不是一般学习问题预先给定的输入。

$$i\neq j,\qquad u_i=u_j\quad\Longrightarrow\quad\text{two sample copies are attributed to the same individual}$$

一般情形先有 Population selector \(U\)。给定 population event \(\mathcal O\),world side 直接形成 \(P(U\in du\mid\mathcal O)\),learner 再以 \(Q_\phi(du\mid\mathcal O)\) 近似;不是 evidence 执行 selection。sample index 只定位 realized value,不能充当 unit index。

$$\mathcal O=\{X=x_i\}\ \text{or}\ \{X^{\mathrm{hist}}=x_i,Y^{\mathrm{hist}}=y_i\},\qquad Q_\phi(du\mid\mathcal O)\approx P(U\in du\mid\mathcal O)$$

Random-variable typing:\(X\) 是 Population-level generic variable,\(X_U\) 是先随机选择 \(U\) 后得到的 selected-unit variable,\(X_u\) 是把 unit 固定为 \(u\) 后的 variable;\(X_{u_i}\) 是由 sample \(i\) 的 attributed unit 索引的 fixed-individual stochastic object,而 \(x_i\) 只是与 \(X_{u_i}\) 相联系的 event-level realized value。本文不把 \(X_i\) 定义为 \(X_{u_i}\),event index 与 unit-indexed stochastic object 不是同一对象(A sample is an event; a unit is not a row index.);这一边界也不决定 repeated observations 的 IID 或 joint law。\(\mathcal O\) 是 event / admissible information;designed candidate \(x^q\) 不会因为被提出就成为 evidence,而 answer 前已观察且 admissible 的 factual \(X=x\) 必须进入 \(\mathcal O\)。

Conditioning precision:正概率 event 使用 ordinary conditional;连续 observed value 使用 regular conditional kernel 的一个 version,而不是对零概率 singleton 做比值。\(\mathcal O\) 中若含 \(Y^{\mathrm{hist}}=y_i\),它只能是已观察的历史 response 或 explanation task 中已发生的 factual outcome,不是当前 prediction target \(Y^q\)。Foundation 不再引入 indexed selector copies;repeated observations 的 joint assignment / sampling law 必须另行声明。

需要抽样时怎样写:使用 evidence-explicit 的 \(\widetilde U_\phi\mid\mathcal O\sim Q_\phi(\cdot\mid\mathcal O)\)。它是 learner-side 抽到的 candidate unit,不是 belief measure 本身、world-side 已实现的 individual、sample attribution,也不是 direct-ID 模型在固定地址学习的 \(u_\theta(k)\)。旧式 \(U_i^{\rm abd}\) 至多是这个 draw 的 deprecated shorthand;因为 \(i\) 不说明 \(\mathcal O\) 是 \(X=x_i\)、\((X,Y)=(x_i,y_i)\)、history 还是 ID event,本页不把 \(U_i\) 作为基础记号。

02 · two sources of randomness

先问是哪一个 individual,再问这个 individual 会发生什么

SOURCE A

Which-unit uncertainty

回答“Population 中是哪一个 individual?”

  • \(U\sim\Pi\)
  • evidence 诱导 \(P(U\in du\mid\mathcal O)\)
  • learner 用 \(Q_\phi\) 逼近
SOURCE B

Fixed-individual event noise

回答“已经固定这个 individual,response 还会怎样随机?”

  • \(U=u\) 固定主体
  • world kernel 保留 exogenous variation;\(P_\theta(dy\mid x^q;u)\) 是 learner approximation
  • 换 query 不等于换 individual
$$U\sim\Pi,\qquad Q_\phi(du\mid\mathcal O)\approx P(U\in du\mid\mathcal O),\qquad P_{\theta,\phi}(dy\mid x^q,\mathcal O)=\int_{\mathcal U}P_\theta(dy\mid x^q;u)Q_\phi(du\mid\mathcal O)$$

严格边界:epistemic uncertainty 位于 learner 对 selected individual 的 belief;event / exogenous noise 位于 fixed-\(u\) response law。把两者压成一个 latent posterior 会失去论文要澄清的类型结构。

POPULATION → INDIVIDUAL
\(U\sim\Pi\)selection:Population 中是哪一个?
\(U=u\)response:fixed referent 语义索引 response law

sample attribution、ID resolver 与 learned response content \(u_\theta(k)\) 都必须接在这条主链的明确位置,而不能替代 selector 或 persistent referent。

03 · key, address, content

已知 unit,为什么 response content 仍然需要学习?

trusted ID kresolver 的输入
\(k\mapsto u(k)\)persistent referent + 稳定 response/update address
\(u_\theta(k)\)该地址上学习的 response content

Direct-ID implementation 需要分清四个角色:\(k\) 是 trusted ID,也是 resolver 的输入;\(u(k)\) 是由 \(k\) 确定的 persistent referent 与稳定 response/update address;\(u_\theta(k)\) 是在该地址由 response objective 学习的 embedding、preference factor、random effect 或其他 numerical parameter block;history-dependent response state 会随时间改变、但不改变 \(u(k)\)。这里不引入 foundational 两对象 ontology;\(u_\theta(k)\) 只是固定 unit response law 的一种 task-specific parameterization,属于 \(\theta\)。多个 records 因为 resolver 指向同一地址,才共同更新这一内容。

\(u_\theta(k)\) 可以与下游 response map 联合旋转、置换或重参数化而不改变预测,所以其坐标通常不唯一;学习它不等于恢复了 person 本身。Token 场景同样如此:token occurrence 提供 realized evidence;token ID 可以揭示 known token-type referent 并确定稳定 parameter address;embedding row 是该地址上的 learned response content。只有 declared attribution protocol 才能说多个 occurrences 属于同一 individual/unit;embedding vector equality 本身不定义这种 identity。

$$K=k,\qquad P(Z_k\in dz\mid\text{ratings},K=k)$$

PMF 的典型结构中 referent 已知,右式是该 known referent 的 preference-property posterior law。它属于 response/property layer,不是 \(P(U\in du\mid\mathcal O)\) 所回答的 which-individual 问题。推荐日志若可靠记录 account identity,equalities 如 \(u_1=u_3=u_A\) 由 known attribution 保证,表达的是 referential equality,不是 learned embedding 的坐标相等。

Dirac boundary:可靠 ID 只有在 protocol 明确说它揭示 selected individual 时,才使 which-individual conditional 退化为 \(\delta_{u(k)}\)。这个 Dirac 在 persistent referent 上,绝不在 trainable \(u_\theta(k)\) 上;embedding 或 random-effect 的估计误差仍属于 response model。

完整类比在推荐系统:user 是 unit,interaction 是 sample,user_id 是 resolver key,item / slate 是 treatment,feedback / retention 是 outcome。

04 · where the analogy breaks

Known referent、稳定地址、learned content 与 learner belief 必须先区分角色

CLOSED → OPEN WORLD

referent 未必由 registry 揭示

known account 可以有可信 address;匿名 session、新患者或新设备需要由 evidence 约束 which-unit selection。

WORLD ≠ LEARNER

\(P\) 与 \(Q_\phi\) 不能互换

\(P(U\in du\mid\mathcal O)\) 是 world-side conditional;\(Q_\phi\) 是 learner-specific epistemic belief。

ADDRESS ≠ CONTENT

\(u(k)\) 与 \(u_\theta(k)\) 是两层

resolver key 固定 persistent referent 与稳定 update address;该地址上的 embedding / factor 由 response objective 学习,坐标可重参数化。

UNIT ≠ EPISODE STATE

持续对象可以变化

同一个 unit 的 context、event noise 与 candidate-conditioned state 都可以随 episode 改变。

PROPERTY ≠ IDENTITY

参数相同不等于同一个人

不同 individuals 可以在当前模型中具有相同 latent property 或 response law。

FIXED U ≠ DETERMINISTIC Y

固定主体仍可随机响应

选择 uncertainty 消失后,event / exogenous noise 仍由 fixed-individual response law 表达。

安全表述:trusted ID → known referent + stable addressaddress → learned content → response 是两条相邻但不同的链。只有显式 identification protocol 才能让 known ID collapse world-side which-unit conditional,且 collapse 发生在 referent 上;embedding lookup 或 property posterior 本身做不到。

理论折叠:evidence reuse、trusted-ID protocol 与两类坐标不唯一性
05 · persistence and query reuse

固定 factual evidence 与一份 learner belief,多个 response queries 复用它

$$Q_\phi(du\mid\mathcal O^F)\ \text{formed once},\qquad (x^Q,c^Q)\ \text{varies}\quad\Longrightarrow\quad\left\{\int_{\mathcal U}P_\theta(\,\cdot\mid x^Q,c^Q,u)Q_\phi(du\mid\mathcal O^F)\right\}$$

query 可以改变,模型也可以做 candidate-conditioned computation;禁止的是 candidate 重新定义 world-side unit、重算 answer-time belief,或把回答之后才出现的 outcome 倒灌进 evidence。新 evidence 改变的是 conditional information 与 learner belief,不会改写已经发生的 world-side selection。

记号边界:\(\mathcal O^F\) 是 answer-time 可用的 admissible factual information;若 protocol 已可靠观察 ID / attribution,它可以包含该信息,但不因此改写成 unknown-\(u\)-indexed evidence。fixed-\(u\) 后直接改变 response 的已知信息显式声明为 \(c^Q\);整包 \(\mathcal O^F\) 不回灌 response kernel。

06 · trusted-ID boundary

只有 identification protocol 才能 collapse which-unit uncertainty

01

Trusted referent

若 registry protocol 保证 \(K=k\) 唯一揭示 persistent referent \(u(k)\),which-individual conditional 才成为 Dirac;identity law 可以是 \(\delta_{u(k)}\),但绝不形成 embedding-row Dirac。

$$P(U\in du\mid K=k)=\delta_{u(k)}(du),\qquad Q_\phi(du\mid K=k)\approx\delta_{u(k)}(du)$$
02

Ambiguous evidence

sparse / ambiguous / new unit 下,world conditional 与 learner belief 都可 non-degenerate,也可以是 population fallback 或 abstention,再对 fixed-individual response law marginalize。

$$P_{\theta,\phi}(dy\mid x,\mathcal O)=\int_{\mathcal U}P_\theta(dy\mid x;u)Q_\phi(du\mid\mathcal O)$$

关键因果顺序是 protocol identifies individual → world conditional collapses → learner may approximate that collapse。ID 字段、embedding row 或模型 posterior 本身都不能倒过来证明 referent 已被识别;deterministic point summary(mean / MAP)至多是算法需要时的 approximation,不是 Unit Abduction 的基础输出。

07 · two kinds of non-uniqueness

Referential relabeling 与 representation gauge 不能混为一谈

I₁

Referential relabeling

对所有 persistent referents 做一致的双射 relabeling \(T\),同时重写 response indexing,模型可以保持相同预测;这不改变谁与谁是同一个 unit。

$$u'=T(u),\qquad f'(x;T(u))=f(x;u),\qquad u_i=u_j\iff T(u_i)=T(u_j)$$
I₂

Representation gauge

另一个不同的不唯一性发生在 response layer:可以联合变换 \(u_\theta(k)\) 与 downstream map 而保持 composed prediction 不变;learned coordinates 不是 person 的唯一科学坐标。

前者是 unit domain 的 referential relabeling;后者发生在 response layer。两者都不把 Unit Abduction 与 embedding estimation 混成同一问题。最后的 equality relation 在 observed-attribution regime 的任意一个固定参数化内部保持。

08 · division of labor

这页负责 typing;相邻两页负责后果与完整案例

01

为什么 unit / sample 必须分开?

计数、加权、split、known/new-unit generalization 与 sample-only boundary 由基础页完整说明。

03

Known referent 与 latent property 怎样进入真实模型?

账户 42、interaction log、trusted ID、user property、item treatment、retention outcome 与 PMF response-layer boundary 由推荐案例完整跑通。

09 · failure modes

六种最常见的类型错误

把未知的 \(u_i\) 预填成每条 sample 的一般输入
把 world conditional \(P\) 与 learner belief \(Q_\phi\) 混写
把 ID lookup 或 embedding posterior 当成 unit selector
仅因出现 user_id 就宣称 learner belief 自动 Dirac
把 which-unit uncertainty 与 fixed-u event noise 混成一个随机源
让 candidate 或 post-answer outcome 偷渡进 evidence
10 · implications for USL-01

这页把什么固定进 Unit-as-Primitive 主线

Bayes rule、latent embedding、random effects 与 posterior predictive 都不是这里的新零件。USL-01 的理论责任是:

在 machine learning 中显式引入 Population selector \(U\),并把 selected individual、sample attribution、world conditional、learner belief、fixed-individual response 与 event-typed evidence 写成可检查的结构。

CHECK 01

Population / Individual

\(U\sim\Pi\) 与 event \(U=u\) 必须在类型上分开。

CHECK 02

World / learner

\(P(U\in du\mid\mathcal O)\) 与 \(Q_\phi\) 必须明确区分。

CHECK 03

Two random sources

Population \(U\) 的 which-individual variation、learner \(Q_\phi\) 的 epistemic uncertainty 与 fixed-u event noise 分开。

CHECK 04

Observed attribution

\(u_i\) 只在 rows 已可靠归属时作为 special regime 使用。

CHECK 05

Referent / address / content

\(u\) 是 persistent referent;direct-ID 下 \(k\mapsto u(k)\) 固定稳定地址,\(u_\theta(k)\) 由 response objective 在该地址学习、可重参数化。

这些是 paper-level theory checks,不是已经通过独立复核的 theorem 或 empirical result。

Owner directive applied

这一页现在固定什么

  1. \(U\) 是执行 individual selection 的 Population Unit Selection Variable,\(U=u\) 固定/表示 definite realized individual;\(\Pi\) 不直接规定 observed dataset sampling。
  2. fixed-\(u\) kernel 保留 event / exogenous variation,且不自动假设 independence。
  3. realized values 用于指定/编码 observed information \(\mathcal O\),其 admissibility 由 protocol 决定。
  4. \(P(U\in du\mid\mathcal O)\) 是 world conditional,\(Q_\phi(du\mid\mathcal O)\) 是 learner belief;factual evidence \(\mathcal O^F\) 一次形成 belief,多个 response queries 复用它,query 不重新进入 Unit Abduction。
  5. indexed \(U_i/U_k\) 是 deprecated ambiguity,不是 foundation object。
  6. \((x_i,y_i;u_i)\)、\(i\mapsto u_i\) 与 \(u_i=u_j\) 仅属于 observed-attribution special regime。
  7. \(u\) 是 task-declared persistent referent 并语义索引 fixed-unit response law;direct-ID 下 \(k\mapsto u(k)\) 固定地址,而 \(u_\theta(k)\) 由 response objective 学习、可重参数化,且其 property posterior 不等同于 which-individual conditional。
  8. fixed-\(u\) 后直接改变 response 的信息必须声明为 \(c^Q\) 并进入 \(P_\theta(dy^Q\mid x^Q,c^Q,u)\);admissible factual \(X=x\) 应进入 \(\mathcal O\),designed candidate \(x^q\) 不会因为进入 response law 就自动成为 \(\mathcal O\);整包 \(\mathcal O^F\) 不回灌 response kernel。
  9. \(\mathcal O\) 中若含 \(Y=y\),它只能是已观察历史 / factual evidence,而不是当前 prediction target \(Y^q\)。
Primary anchors

Embedding 与推荐 response-layer 邻居的一手来源

这些来源分别支撑 token machinery 与已有推荐实践,不支撑 Unit-Abductive Learning 的 novelty。unit ontology、belief reuse、invariance、identification 与 empirical advantage 仍需单独证明与审计。