USLTen-paper research program
USL-01 · Research project · Foundations

The Unit as a Primitive for Machine Learning: A Formal Foundation for Learning from Unit-Linked Events

这个项目把跨 events 持续存在的 task-declared unit 提升为机器学习的语义 primitive,并研究它如何组织 response-law family、shared structure 与 learner access。

01 · The project question

先声明谁跨 events 持续存在,再决定围绕它学习什么

USL-01 的根命题是 unit primitive;heterogeneity、learner access 和因果 specialization 都是从这个根向下展开的研究层。

What changes when a learning task first declares which persistent referents count as units and how events are attributed, before choosing an observational or causal learning object; how is this developed through supervised response laws and instantiated separately by DiscoSCM under additional causal commitments?

Persistent referent

Unit 不是 sample index

sample 是一次 observed event;unit 是由当前 task 声明、跨相关 events 保持 referential identity 的对象。

Learned object

一族 response laws

监督学习的主要 specialization 把单一总体关系展开为共享结构下的 unit-conditioned response-law family。

Learner access

两条访问路线

trusted resolver 直接供应 unit;只有没有 resolver 时,factual evidence 才形成可跨 queries 复用的 which-unit belief。

核心 claim
A learning task first declares which persistent referents count as units and how events remain linked to them, then chooses the task-specific learning object; supervised response laws are the main formal specialization developed here, while DiscoSCM provides a distinct causal specialization under additional commitments and unit abduction appears only when no resolver supplies the realized unit.
如果成立会改变什么
Population-level which-individual variation, learner-specific epistemic uncertainty, and within-individual event variation become distinct mathematical objects, clarifying what a sample records, when attribution is known, what evidence can update unit belief, how prediction propagates both sources of randomness, and which generalization or split question is actually being asked.
02 · Project map

项目首页说明全局关系,技术细节回到各自的完整 home

Foundations 先建立 unit / sample 与类型边界;Theory Article 展开 response heterogeneity;Related Materials 承担 prior-art pressure。三者不在入口重复全文。

Theory article · owner review

异质的是谁?

在 unit primitive 之后,展开 unit-conditioned response heterogeneity、shared structure 与 historical map。

03 · Current paper

Unit primitive 是唯一根;监督学习是主要 formal specialization

A learning task first declares which persistent referents count as units and how events remain linked to them, then chooses the task-specific learning object; supervised response laws are the main formal specialization developed here, while DiscoSCM provides a distinct causal specialization under additional commitments and unit abduction appears only when no resolver supplies the realized unit.

查看已登记的 OpenReview record ↗
Public semantics
v0.12 · noindex owner-review projection
Local working source
v0.15 · mutable 43-page build · targeted visual QA only
Accepted artifacts
Main v0.14 · supplement v0.9
External boundary
Registered title remains separately tracked; the current local title, abstract and source are not claimed synchronized.
展开 manuscript spine、所需证据与 working abstract

论证 spine

  1. Introduction
  2. The Unit as a Machine-Learning Primitive
  3. Learning a Family of Unit-Conditioned Response Laws
  4. Learner Access to the Unit
  5. Unit Information: Predictive Value, Access, and Deployment
  6. Recommendation as a Worked Setting
  7. Relations to Established Learning Traditions
  8. Empirical Implications and Evaluation
  9. Discussion and Limitations
  10. Conclusion

所需证据

The unit layer is scientifically useful only if preserving it changes a declared question, estimand, evaluation protocol, or falsifiable prediction. Proposed tests must separate unit-belief quality, fixed-unit response quality, marginalized prediction, and a matched unit-omitting baseline. They must use evidence/target separation, unit-disjoint evaluation where unseen-unit claims are made, unit-abduction settings where relevant, and negative controls such as grouping shuffles, label permutations, and no-persistence conditions.

No experiment has yet been run. If the unit-preserving formulation cannot be distinguished from an equally informed unit-omitting formulation on a requested or testable quantity, the simpler abstraction should be used.

Working abstract(未声称已在外部登记)

Machine learning is usually formalized through samples, while the persistent individual to which multiple observed or possible events refer often remains implicit. We propose the *unit* as an explicit primitive at the level of task semantics. A learning task first declares a population of persistent referents and a criterion for when observed or possible events concern the same one; the realized value $u$ denotes the selected referent. We develop supervised learning as the main formal specialization. Its learned object is a structured family of unit-conditioned response laws: homogeneity is the special case in which the laws coincide across units, while heterogeneity allows them to differ under shared restrictions. A sample-only conditional law is silent about this distinction because it may be either a homogeneous unit extension or the marginal projection of a heterogeneous family. Learning therefore requires separate statements about what varies across units and what structure permits information sharing among them. For many applied tasks, this makes stable individual variation available as learnable structure and inductive bias rather than leaving it only as residual error; it does not presume that every task is heterogeneous. Learner access is downstream of this declaration. A trusted resolver may provide direct unit access; when no resolver identifies the realized unit, *unit abduction* maps factual evidence to a belief over that already-realized unit and reuses the belief across response queries. Under response sufficiency, we separate oracle predictive value, value accessible from declared evidence, learned-pipeline approximation, and component identification. We also show that arbitrarily many unlinked single-row observations can fail to distinguish a heterogeneous unit world from a homogeneous pooled world, whereas trusted same-unit pairs separate a restricted witness. Practical corollaries clarify weighting and data splitting. The formulation organizes established repeated-measures, personalization, entity-resolution, latent-variable, and heterogeneous-model traditions around one unit contract without claiming their machinery as new. Its formal results concern the supervised predictive specialization; observational, causal, and other learning objects require their own additional commitments.

04 · Current state

已经固定研究契约,但还没有把方向写成已验证结果

Fixed at contract level

当前已经明确

semantic unit、sample/event 区分、structured response family、direct access / unit abduction 两条路线,以及现有 theorem spine。

Open evidence gates

仍待完成

独立数学与文献复核、proper-score promotion 判断、真实实验、population geometry、freeze 与 external readback。

Evidence boundary

当前不能声称

经验优势、unique unit recovery、positive causal identification、穷尽 novelty 判断,或当前本地稿已经同步外部记录。

下一关

先独立复核当前 unit-primitive 逻辑链与 theorem boundary,再运行能区分 unit-preserving 和 unit-erased formulation 的 falsification protocol。

展开 exact status、evidence gates 与 provenance

Exact current status

public semantic contract v0.12 is preserved. On 2026-07-30, the local canonical v0.15 source was grown so that the first-reading layer now explicitly answers which default machine-learning abstractions the unit primitive revises: a sample is an event rather than the persistent semantic subject; the unit precedes IDs and learned representations; a population may carry a structured family of unit-conditioned response laws rather than only one subject-free map; structured heterogeneity is a first-class model possibility for many applied tasks, with homogeneity retained as a special case; and which-unit attribution is separated from response inference. It also distinguishes referential identification from response-relevant structural localization, states the conditional duality between localization and information borrowing, and presents cross-field partial recovery as a research program rather than a reduction theorem. The existing value--access--approximation--identification theorem spine remains unchanged. The mutable 43-page working build compiles without manuscript warnings and has passed targeted visual inspection of the changed first-reading and conclusion surfaces, but it has not replaced the preserved 2026-07-28 all-page-reviewed preview and is not a frozen artifact. Experiments, theorem reprioritization, exact borrowing geometry, exhaustive historical-priority review, independent external review, freeze, and external synchronization remain open. The frozen v0.14 manuscript and supplement v0.9 remain the latest accepted artifacts.

Next writing gate

independently review the current unit-primitive logical chain, the value--access decomposition, the learned-pipeline component-error propagation, the single-row impossibility and repeated-linkage separation, the direct-access address/content boundary, the historical and DiscoSCM attribution boundaries, and the weighting/split diagnostics; then run the discriminating unit-preserving versus unit-erased evaluation. The 41-page direct-access/address-content PDF and same-basename SyncTeX artifact have been generated and internally verified (document QA plus focused internal semantic and mathematical audit), but remain unfrozen owner-review artifacts. Exhaustive priority and independent novelty review remain open. Do not treat working-source document QA, the local AAAI adaptation, or the frozen v0.14 artifact as public or externally synchronized.

Next evidence gate

freeze and run whole-user cold-start recommendation plus one non-isomorphic repeated-measure task with negative controls.

48h minimal experiment

implement a whole-user cold-start recommender split with $m\in\{0,1,3,5\}$ admissible prior interactions and a small synthetic negative control; compare sample-only prediction, observed-key PMF/property inference, evidence-abducted selector beliefs, deterministic point-summary ablations, Neural Process / Conditional Neural Process context inference, random-effects/grouped and established recommender baselines under matched information.

Closest SOTA opponents

entity/user/token embeddings; grouped latent-variable and Neural Process / amortized-inference models; hierarchical Bayes and random-effects models.

Strongest objection / kill signal

This is only entity embedding, hierarchical latent-variable learning, random effects or grouped representation learning with renamed notation.

downgrade the work to a perspective or tutorial if the Population/Individual distinction, the direct world conditional, and the two-source randomness contract add no falsifiable constraint, evaluation axis or algorithmic operation beyond established grouped or latent-variable practice.

Canonical provenance

highest-priority owner contract
papers/USL01-unit-as-a-primitive-for-machine-learning/hello_agent.md
v0.12 · sha256:1bd6a67c038b
local canonical source: v0.15
latest verified PDF: v0.14
latest verified supplement: v0.9
synchronized mirrors:
research-questions/USL01-unit-as-a-primitive-for-machine-learning/seed.md
papers/USL01-unit-as-a-primitive-for-machine-learning/paper.md

Registration boundary. This is the owner-approved working title. The external AAAI-27 record is still tracked under The Unit as a Primitive for Machine Learning: Foundations of Unit-Abductive Learning; a title update has not been verified.

Progress log

2026-07-29 · the direct-access/address-content round is now canonical. A trusted resolver key k fixes the persistent referent and stable update address u(k), while the response objective may still learn task-specific numerical content uθ(k) at that address; the identity law is Dirac on the persistent referent, never on the trainable representation. The formal spine is value--access decomposition, learned-pipeline realization with component-error propagation, and single-row impossibility with repeated-linkage separation.
2026-07-28 · the current 41-page owner-review preview passed an isolated clean build, bibliography verification, zero-warning and duplicate-label scans, all-page visual QA, a SyncTeX source-navigation check, and focused internal semantic and mathematical audits with no blocking issue. It remains unfrozen, independently unreviewed and externally unsynchronized.
2026-07-28 · local canonical source v0.15 preserves public semantic contract v0.12 while rearchitecting the mainline around the unit primitive. Minimal response heterogeneity, formal consequences, historical synthesis and falsifiable diagnostics now appear in the body; detailed qualifications, proofs and source-specific maps remain in the appendix.
2026-07-28 · the preceding 35-page learner-access snapshot is retained as provenance. It established the two learner-access routes: direct unit access through a trusted resolver k → u(k), and Unit Abduction forming Qφ(du | OF) only when no resolver supplies the realized unit; the current round sharpens the direct-access route into the address/content distinction.
2026-07-26 · factual evidence OF forms one learner belief Qφ(du | OF). Alternative xQ queries reuse that belief and do not re-enter Unit Abduction.
2026-07-26 · information with a direct response role is typed separately as cQ and enters Pθ(dyQ | xQ, cQ, u); the full abductive evidence bundle is not silently fed back into the response kernel.
2026-07-26 · main PDF v0.14 (22 pages) and supplement v0.9 (40 pages) remain the latest locally verified and frozen artifacts. The v0.15 working source passed document QA but is not frozen.
2026-07-28 · the Chinese heterogeneity article is a noindex WeHub public owner-review projection of one main-text logic and its appendix details. It is not a manuscript/PDF deploy, theorem verification, novelty judgment, formal release, or external-record synchronization.
05 · Reading paths

根据你现在的问题,选择最短阅读路径

01

我想先懂 foundation

按三页路径读 Unit / sample、U=u 与 recommender worked example。

进入 Foundations →

02

我想检验一条理论逻辑

直接进入 unit-conditioned response heterogeneity 长文,查看 historical map 与 non-reduction boundary。

阅读 Theory Article →

03

我想审 prior art

从 Related Materials 进入单项 dossier,比较 hierarchical Bayes / mixed-effects 与 USL-01 的重叠和断点。

进入 Related Materials →

06 · Feedback loop

网页负责阅读,反馈回到 live grow place