Unit 不是 sample index
sample 是一次 observed event;unit 是由当前 task 声明、跨相关 events 保持 referential identity 的对象。
这个项目把跨 events 持续存在的 task-declared unit 提升为机器学习的语义 primitive,并研究它如何组织 response-law family、shared structure 与 learner access。
USL-01 的根命题是 unit primitive;heterogeneity、learner access 和因果 specialization 都是从这个根向下展开的研究层。
What changes when a learning task first declares which persistent referents count as units and how events are attributed, before choosing an observational or causal learning object; how is this developed through supervised response laws and instantiated separately by DiscoSCM under additional causal commitments?
sample 是一次 observed event;unit 是由当前 task 声明、跨相关 events 保持 referential identity 的对象。
监督学习的主要 specialization 把单一总体关系展开为共享结构下的 unit-conditioned response-law family。
trusted resolver 直接供应 unit;只有没有 resolver 时,factual evidence 才形成可跨 queries 复用的 which-unit belief。
Foundations 先建立 unit / sample 与类型边界;Theory Article 展开 response heterogeneity;Related Materials 承担 prior-art pressure。三者不在入口重复全文。
先看 persistent unit 与 observed event 的基本区别,再进入 U=u 的严格类型页和账户 42 的推荐系统 worked example。
查看当前 claim、working source、latest accepted artifact、未完成证据门和折叠的 manuscript details。
在 unit primitive 之后,展开 unit-conditioned response heterogeneity、shared structure 与 historical map。
公网当前提供 RM01;RM02 / RM03 仍只保存在 canonical materials,等待独立投影。
A learning task first declares which persistent referents count as units and how events remain linked to them, then chooses the task-specific learning object; supervised response laws are the main formal specialization developed here, while DiscoSCM provides a distinct causal specialization under additional commitments and unit abduction appears only when no resolver supplies the realized unit.
查看已登记的 OpenReview record ↗The unit layer is scientifically useful only if preserving it changes a declared question, estimand, evaluation protocol, or falsifiable prediction. Proposed tests must separate unit-belief quality, fixed-unit response quality, marginalized prediction, and a matched unit-omitting baseline. They must use evidence/target separation, unit-disjoint evaluation where unseen-unit claims are made, unit-abduction settings where relevant, and negative controls such as grouping shuffles, label permutations, and no-persistence conditions.
No experiment has yet been run. If the unit-preserving formulation cannot be distinguished from an equally informed unit-omitting formulation on a requested or testable quantity, the simpler abstraction should be used.
Machine learning is usually formalized through samples, while the persistent individual to which multiple observed or possible events refer often remains implicit. We propose the *unit* as an explicit primitive at the level of task semantics. A learning task first declares a population of persistent referents and a criterion for when observed or possible events concern the same one; the realized value $u$ denotes the selected referent. We develop supervised learning as the main formal specialization. Its learned object is a structured family of unit-conditioned response laws: homogeneity is the special case in which the laws coincide across units, while heterogeneity allows them to differ under shared restrictions. A sample-only conditional law is silent about this distinction because it may be either a homogeneous unit extension or the marginal projection of a heterogeneous family. Learning therefore requires separate statements about what varies across units and what structure permits information sharing among them. For many applied tasks, this makes stable individual variation available as learnable structure and inductive bias rather than leaving it only as residual error; it does not presume that every task is heterogeneous. Learner access is downstream of this declaration. A trusted resolver may provide direct unit access; when no resolver identifies the realized unit, *unit abduction* maps factual evidence to a belief over that already-realized unit and reuses the belief across response queries. Under response sufficiency, we separate oracle predictive value, value accessible from declared evidence, learned-pipeline approximation, and component identification. We also show that arbitrarily many unlinked single-row observations can fail to distinguish a heterogeneous unit world from a homogeneous pooled world, whereas trusted same-unit pairs separate a restricted witness. Practical corollaries clarify weighting and data splitting. The formulation organizes established repeated-measures, personalization, entity-resolution, latent-variable, and heterogeneous-model traditions around one unit contract without claiming their machinery as new. Its formal results concern the supervised predictive specialization; observational, causal, and other learning objects require their own additional commitments.
semantic unit、sample/event 区分、structured response family、direct access / unit abduction 两条路线,以及现有 theorem spine。
独立数学与文献复核、proper-score promotion 判断、真实实验、population geometry、freeze 与 external readback。
经验优势、unique unit recovery、positive causal identification、穷尽 novelty 判断,或当前本地稿已经同步外部记录。
public semantic contract v0.12 is preserved. On 2026-07-30, the local canonical v0.15 source was grown so that the first-reading layer now explicitly answers which default machine-learning abstractions the unit primitive revises: a sample is an event rather than the persistent semantic subject; the unit precedes IDs and learned representations; a population may carry a structured family of unit-conditioned response laws rather than only one subject-free map; structured heterogeneity is a first-class model possibility for many applied tasks, with homogeneity retained as a special case; and which-unit attribution is separated from response inference. It also distinguishes referential identification from response-relevant structural localization, states the conditional duality between localization and information borrowing, and presents cross-field partial recovery as a research program rather than a reduction theorem. The existing value--access--approximation--identification theorem spine remains unchanged. The mutable 43-page working build compiles without manuscript warnings and has passed targeted visual inspection of the changed first-reading and conclusion surfaces, but it has not replaced the preserved 2026-07-28 all-page-reviewed preview and is not a frozen artifact. Experiments, theorem reprioritization, exact borrowing geometry, exhaustive historical-priority review, independent external review, freeze, and external synchronization remain open. The frozen v0.14 manuscript and supplement v0.9 remain the latest accepted artifacts.
independently review the current unit-primitive logical chain, the value--access decomposition, the learned-pipeline component-error propagation, the single-row impossibility and repeated-linkage separation, the direct-access address/content boundary, the historical and DiscoSCM attribution boundaries, and the weighting/split diagnostics; then run the discriminating unit-preserving versus unit-erased evaluation. The 41-page direct-access/address-content PDF and same-basename SyncTeX artifact have been generated and internally verified (document QA plus focused internal semantic and mathematical audit), but remain unfrozen owner-review artifacts. Exhaustive priority and independent novelty review remain open. Do not treat working-source document QA, the local AAAI adaptation, or the frozen v0.14 artifact as public or externally synchronized.
freeze and run whole-user cold-start recommendation plus one non-isomorphic repeated-measure task with negative controls.
implement a whole-user cold-start recommender split with $m\in\{0,1,3,5\}$ admissible prior interactions and a small synthetic negative control; compare sample-only prediction, observed-key PMF/property inference, evidence-abducted selector beliefs, deterministic point-summary ablations, Neural Process / Conditional Neural Process context inference, random-effects/grouped and established recommender baselines under matched information.
entity/user/token embeddings; grouped latent-variable and Neural Process / amortized-inference models; hierarchical Bayes and random-effects models.
This is only entity embedding, hierarchical latent-variable learning, random effects or grouped representation learning with renamed notation.
downgrade the work to a perspective or tutorial if the Population/Individual distinction, the direct world conditional, and the two-source randomness contract add no falsifiable constraint, evaluation axis or algorithmic operation beyond established grouped or latent-variable practice.
highest-priority owner contractpapers/USL01-unit-as-a-primitive-for-machine-learning/hello_agent.mdv0.12 · sha256:1bd6a67c038b
local canonical source: v0.15
latest verified PDF: v0.14
latest verified supplement: v0.9
synchronized mirrors:research-questions/USL01-unit-as-a-primitive-for-machine-learning/seed.mdpapers/USL01-unit-as-a-primitive-for-machine-learning/paper.md
Registration boundary. This is the owner-approved working title. The external AAAI-27 record is still tracked under The Unit as a Primitive for Machine Learning: Foundations of Unit-Abductive Learning; a title update has not been verified.
k fixes the persistent referent and stable update address u(k), while the response objective may still learn task-specific numerical content uθ(k) at that address; the identity law is Dirac on the persistent referent, never on the trainable representation. The formal spine is value--access decomposition, learned-pipeline realization with component-error propagation, and single-row impossibility with repeated-linkage separation.v0.15 preserves public semantic contract v0.12 while rearchitecting the mainline around the unit primitive. Minimal response heterogeneity, formal consequences, historical synthesis and falsifiable diagnostics now appear in the body; detailed qualifications, proofs and source-specific maps remain in the appendix.k → u(k), and Unit Abduction forming Qφ(du | OF) only when no resolver supplies the realized unit; the current round sharpens the direct-access route into the address/content distinction.OF forms one learner belief Qφ(du | OF). Alternative xQ queries reuse that belief and do not re-enter Unit Abduction.cQ and enters Pθ(dyQ | xQ, cQ, u); the full abductive evidence bundle is not silently fed back into the response kernel.v0.14 (22 pages) and supplement v0.9 (40 pages) remain the latest locally verified and frozen artifacts. The v0.15 working source passed document QA but is not frozen.按三页路径读 Unit / sample、U=u 与 recommender worked example。
直接进入 unit-conditioned response heterogeneity 长文,查看 historical map 与 non-reduction boundary。
从 Related Materials 进入单项 dossier,比较 hierarchical Bayes / mixed-effects 与 USL-01 的重叠和断点。