本页的 \(U,Q_\phi\) 与 “Unit-Abductive” 只保留为 historical v0.3 terminology。当前 executable object 仅是 row/evidence-conditioned latent response-modulation law:没有 Population selector,没有 world conditional 与 learner approximation contract,不支持 \(Q_\phi\approx P(U\mid\mathcal O)\),也没有 repeated-individual linkage 或 exact USL-01 claim。因此不能作为 USL-01 exact selector reduction。Branch A / B 的 owner decision 仍然 pending;本页不预先选择任何分支。
P32 · historical v0.3 computational schema · scientific type unresolved
Unit-Abductive Local Bilinear Models (historical v0.3 title)
当前可执行模型从每条 row 的 evidence 产生 latent-coordinate location/scale law,经 bilinear coefficient modulation 传播到最终预测分布。Cauchy / Gaussian coordinate laws 都可解析边缘化;这不构成“推断哪一个 Population individual”。九数据集 diagnostics 仍只支持窄范围的 gross-label-outlier signal。
P32 现在是 theory + prospective diagnostics 的 factual-prediction preprint:v0.3 protocol 在项目内部冻结后运行,但没有 external preregistration;它仍是固定小样本 screen,不是调参 benchmark,也没有一般性能优势或一般鲁棒性主张。当前页面、英文 PDF 与 3-file source bundle 已作为 public owner-review preview 提供;公开可读不等于正式论文发布,arXiv upload 与 venue submission 仍须单独批准。
1. 一句话理解
当前 v0.3 encoder 从一条 sample row 的 evidence \(O=o\) 直接产生 historical coordinate \(U\) 上的 location/scale law;candidate \(u\) 调制 predictor \(X\) 的 slope 与 intercept。固定 \(u\) 后模型对 \(X\) 是 affine,并通过 \(X^\top BU\) 形成 bilinear interaction。把 stable coordinate law 积分掉以后,\(Y\mid X,O\) 仍有闭式 stable distribution。数据契约没有 individual key 或 repeated-individual linkage,所以这里不能读成 selector ontology。
2. Historical v0.3 latent-coordinate law
在当前 executable contract 中,\(U\in\mathbb R^d\) 只是进入 response head 的 continuous latent modulation coordinate;\(Q_\phi\) 由每条 row 的 evidence 直接参数化。它没有 selector truth、persistent referent、individual key 或 repeated-individual constraint。历史符号写作
当前许可的语言是:“encoder 由 row evidence 产生一个 latent-coordinate law”。不能把它升级为 learner 对 Population individual 的 subjective unit belief,也不能写成对 \(P(U\mid\mathcal O)\) 的近似。
point coordinate 只给一个 location;当前 encoder law 还给 scale,并把二者一起传播到 predictive distribution。如果只把 \(m_\phi(o)\) 输入 MSE predictor,而完全丢掉 \(\gamma_\phi(o)\),就不再是当前 distribution-valued computation。
这个 law 也不自动等于 calibrated posterior 或 epistemic uncertainty。当前 predictive NLL 只约束 composed response law;它没有 selector-calibration target。
3. 为什么斜率、截距关于 \(u\) 线性会得到 bilinear model
模型先定义 coordinate-conditioned slope 与 intercept:
再令
展开以后:
- \(\beta^\top X\) 是共享的 predictor effect;
- \(a^\top U\) 是 latent-coordinate main effect;
- \(X^\top BU\) 让 latent coordinate 改变 \(X\) 的 slope;
- \(\alpha_0\) 是共享截距。
所以 bilinear 不是装饰性的命名,而是 “\(w,b\) 对 \(u\) affine” 的精确代数结果。
4. `local` 的正式含义
固定一个 latent-coordinate candidate \(u\) 后,\(X\mapsto Y\) 是 affine。row evidence 产生的 location/scale law 决定 coordinate-conditioned view。这就是当前 P32 可执行对象中的 locality;它不等于 individual locality。
它不是数据空间中的 top-\(k\)、nearest-neighbor WLS、prototype routing 或 query-time matrix solve。P31 的 “local KRR” 与 P32 v0.3 的 “coordinate-conditioned local view” 必须保持术语边界。
Taylor theorem 提供什么
对在 \((x_0,u_0)\) 附近足够光滑的 \(f(x,u)\),令 \(\Delta x=x-x_0\)、\(\Delta u=u-u_0\),保留常数、一阶 main effects 和中心处 mixed-Hessian term:
未保留部分由 pure-\(x\) curvature、pure-\(u\) curvature 与三阶余项控制。这个 theorem 支持“局部 bilinear 是合理 inductive bias”,不支持“任意非线性都被单个 bilinear model 全局、精确表示”。
理论子页从 \(f_\star(x)=x^2\) 的可移动切线开始,逐步区分真实 tangent coefficient field、latent-coordinate affine approximation、row-conditioned coordinate law,以及 within-view / view-adaptation 两种导数。
打开:为什么先选择线性?→5. Stable latent-coordinate law 的闭式传播
截距已经使用 \(\alpha_0\),因此 stable index 记作 \(\rho\in(0,2]\)。本文用
数学子页用“旋转箭头探针”解释 \(z_{ik}\)、\(c_k\) 和完整 CF matching,同时记录最新方法决定:有限 Cauchy mixture 一般不是 Cauchy,把 source aggregate 强行拉成标准 Cauchy 不适合作为默认训练目标;当前只保留 median/MAD loc-scale 规范。子页还给出 learned effective-slope source density,并专门对照 UALBM 与 VAE。
阅读 learned source density、被排除方案与 UALBM vs VAE →定义 \(S_\rho S(m,\gamma)\)。设给定 row evidence 后,各 latent coordinates 条件独立:
令 \(g(x)=a+B^\top x\),则
“closed-form” 到这里为止:marginal predictive law 与 likelihood 闭式,但 neural coordinate-law map 和 bilinear parameters 仍然通过 numerical maximum likelihood 训练。
Cauchy 主线:\(\rho=1\)
预测 scale 为
闭式 Cauchy NLL 为
Cauchy 的 location score 对大残差有界,但这只是 analytic robustness property,不是经验 superiority。由于 Cauchy 没有有限 mean 与 variance,论文只报告 location、scale 和 central quantiles。
Gaussian 匹配基线:\(\rho=2\)
在相同 characteristic-function convention 下:
因此 Gaussian standard deviation 必须写成 \(\sqrt 2\,s_2\)。这个换算让 Gaussian 与 Cauchy 使用同一个 stable propagation theorem,而不会把 stable scale 错当成标准差。
三种网络架构
三条 lane 的 evidence-to-hidden-head 采用同一种架构语言,但参数各自独立训练。Direct MLP 把 deterministic representation 与 \(X\) 送入 scalar point-output head;它没有 distribution-valued latent coordinate。Gaussian 与 Cauchy 使用同构的 location/positive-scale heads 和 analytic bilinear marginalizer,只改变 \(\rho\)、scale aggregation 与 endpoint NLL。
6. 理论主干与可观测内容
- Centered local-bilinear approximation:给出 mixed term 与 remainder bound。
- Affine-coefficient equivalence:证明 coefficient modulation 与 bilinear representation 等价。
- Stable marginalization:用 characteristic functions 推导一般 \(S_\rho S\) predictive law。
- Special cases:Cauchy / Gaussian NLL、quantiles、score 与 scale conversion。
- Proper-score target:正确指定 family 内的 population log-score Fisher consistency。
- Non-identifiability:latent basis reparameterization 与 coordinate/event scale trade-off。
不同的 latent coordinate systems 可能诱导同一个 predictive law。P32 的可审计对象是 evidence-conditioned varying coefficients、最终 location/scale 与 predictive distribution,而不是对 \(U\) 坐标的唯一解释。
7. 预测诊断 v0.3:从 discovery 到 prospective extension
实验仍先回答实现层问题:三种 neural realizations 能否作为普通 supervised predictors 运行;development labels 被污染时,matched Cauchy full pipeline 是否比 Gaussian full pipeline 少损失 clean-test prediction。它不是经过调参的 benchmark,也不是 general robustness study。
California Housing、Wine Quality、Abalone 三个数据集上的 180-run 小实验发生在 protocol lock 之前,只用于发现现象、修正数值实现和形成 prospective 问题。它们继续保留为 discovery evidence,但不计入九数据集的 dataset-level gate。
- 数据:Diabetes、Concrete、Energy(heating)、Airfoil、Yacht、Auto MPG、QSAR Fish、Protein、Superconductivity;每个数据集最多抽取 \(n=500\),使用 seeds 42–46。
- 切分:每个 seed 固定 60/20/20 random-row train/validation/test;特征标准化只拟合训练集;test inputs 与 labels 始终干净。
- 四个模型:Direct MLP、Gaussian UALBM、Cauchy UALBM,以及固定默认参数
GradientBoostingRegressor。 - 三个 arms:clean;train/validation 各自精确 20% value-changing label shuffle;各自精确 10% labels 加上 balanced \(\pm5s_y\) gross outliers,其中 \(s_y\) 只由干净训练 target 计算。
- 分析单位:先在每个数据集内对五个 paired seeds 取中位数,再把 dataset 作为九个分析单元;45 个 seed pairs 只作配对敏感性描述。
prospective 矩阵是 9 × 5 × 3 × 4,共 540 个输出;与 discovery 的 180 个输出合计 720 rows。protocol 在查看 prospective 结果前已于项目内部冻结,但没有 external preregistration。
实现与数据 oracle 先于结果解读
float64 oracle 覆盖手算 bilinear case、stable analytic propagation、automatic gradients、singleton/batch equivalence 与 independent latent composition;dataset verifier 还锁定 raw/parsed hashes、feature schema、targets、所有 seeds 的 split digests、污染计划与 clean-test hashes。720 个观测输出全部有限。这里的 finite arithmetic 不等于 statistical robustness,也不证明任意输入上的安全外推。
Clean prediction:stable endpoints 没有获胜
以每个 prospective dataset 的五种子 clean-test RMSE 中位数判定 winner,固定默认 GradientBoosting 赢 7/9,Direct MLP 赢 2/9,Gaussian 与 Cauchy 都是 0/9。树模型完成了“强方法大概能预测到什么量级”的快速参照,也直接阻止了 clean predictive-superiority claim。
Gross label outliers:九数据集上一致,但范围很窄
对每个 dataset–seed,同时计算两项:Cauchy 相对 Gaussian 的 clean-to-dirty RMSE degradation contrast,以及 dirty-arm RMSE gap;两者都除以该数据集的 clean-training target SD。负值表示 Cauchy full pipeline 的 clean-test sensitivity 更小。这个单位是 target-SD unit,不是百分比。
| Prospective arm | Dataset-level 双向支持 | Seed-pair 双向支持 | 归一化 macro medians |
|---|---|---|---|
| 10% balanced \(\pm5s_y\) outliers | 9 / 9 | 41 / 45 | degradation −0.160;dirty RMSE −0.218 |
| 20% value-changing shuffle | 4 / 9 | 19 / 45 | 很弱;bootstrap intervals 跨 0 |
gross-outlier arm 的 −0.160 与 −0.218 分别是九个 dataset medians 的 macro median;它们不能读作 16.0% 或 21.8%。9/9 表示五种子聚合后,九个数据集的两个 contrast 都支持 Cauchy;41/45 是更细的 seed-pair 描述。相比之下,shuffle 只有 4/9 与 19/45,而且 bootstrap intervals 跨过零,所以没有一般 label-noise robustness evidence。
Concrete、QSAR Fish、Protein 与 Superconductivity 在至少一个 seed 中出现跨 train/test 的 exact-\(X\) 重复;其中一部分相同 \(X\) 对应不同 \(Y\)。这不是已经证实的 duplicate-\((X,Y)\) target leakage,但说明 v0.3 只是 random-row screen,不能包装成 unseen-unit 或严格去重后的 generalization test。
下一 evidence gate 是 grouped/deduplicated split、gross-outlier severity curve,以及 loss-only / clean-anchored target-transform ablations。由于 Gaussian/Cauchy full pipelines 同时在 endpoint NLL、stable-scale aggregation 和 dirty-development selection 中存在耦合,v0.3 不能把差异单独归因于 bounded Cauchy score。
8. 当前能说与不能说
| 主题 | 当前可以说 | 当前不能说 |
|---|---|---|
| Historical \(Q_\phi\) | row/evidence-conditioned latent response-modulation law,location/scale 进入预测 | Population selector、learner unit belief、真实 individual 或 calibrated posterior |
| Locality | coordinate-conditioned affine view,受 Taylor remainder 控制 | individual-local semantics,或任意全局非线性精确等于单个 bilinear model |
| Closed form | stable marginal predictive law 与 likelihood | network parameters 存在 analytic MLE |
| Cauchy 理论 | location score 有界,location/scale/quantile 可计算 | mean、variance或把 scale 称为标准差 |
| v0.3 实验 | gross outliers 9/9 datasets、41/45 pairs;shuffle 4/9、19/45;clean winners 是 tree 7/9、MLP 2/9 | 一般鲁棒性、一般预测优越性或把差异只归因于 score |
| 任务 | factual prediction | 干预或因果识别 |
| 状态 | public owner-review theory + prospective diagnostics v0.3 preview | 已正式发布、已投稿、已完成 benchmark 或 claims 已被 owner 接受 |
9. 常见混淆
当前 coordinate law 和 point embedding 有什么不同?
point embedding 通常只保留一个点;当前 encoder law 还保留条件 scale,并把它传播到 predictive distribution。这个计算差异不赋予 selector semantics。
当前 \(Q_\phi\) 是 unit selection 吗?
尚不能这样说。当前 \(Q_\phi(du\mid O_i)\) 只是 row/evidence-conditioned latent-coordinate law;sample selection 则是 rows 之间的 inclusion / weighting 问题。两者都不等于 USL-01 的 Population selector。\(Q_\phi\) 也不是 propensity score,P32 当前 baseline 没有使用 IPW。
打开已暂停教程的 compatibility-hold 说明 →
这是不是换成 Cauchy latent 的 VAE?
不是。VAE 建立 \(p(z)p_\theta(O\mid z)\),用 reconstruction likelihood 与逐样本 prior KL 组成 ELBO;P32 v0.3 直接学习监督式 \(p(Y\mid X,O)\),让 row-conditioned latent coordinate 调制 affine slope/intercept,再解析边缘化为 response law。当前模型没有 \(p(O\mid U)\) decoder、reconstruction 或 ELBO。两者都可使用 amortized distribution-valued encoder;Gaussian / Cauchy 不是根本分界。
为什么没有 top-k 还叫 local?
因为 locality 位于 coordinate-conditioned model view,而不是 reference neighborhood。固定 \(u\) 后 predictor 对 \(X\) 是 affine。
Cauchy scale 能叫 uncertainty variance 吗?
不能。Cauchy variance 不存在;正确对象是 predictive scale 与 central quantile interval。
这是一篇因果论文吗?
不是。\(O,X,Y\) 全部服务于 factual prediction;本文不承诺干预解释、个体效应或 mechanism recovery。
v0.3 已经证明 Cauchy 更鲁棒了吗?
没有。九数据集上 10% gross-outlier arm 的 9/9 dataset-level 与 41/45 seed-pair 双向支持,是这套固定 suite 中针对稀疏大幅标签离群点的窄 signal;20% shuffle 只有 4/9 与 19/45,区间还跨过零。它支持下一轮受控验证,不支持 general robustness claim。
为什么污染 train 和 validation,却保持 test 干净?
因为问题是“dirty development process 会让最终 predictor 损失多少干净目标上的预测能力”。干净 test set 固定了评估目标;若 test labels 也被污染,就无法把训练受损与评估本身变脏区分开。
California seed 45 是 Gaussian 数值溢出吗?
不是。prediction 与 loss 都是 finite。异常来自一个远超训练范围的 predictor,经 \(O=X\) bilinear path 被放大;它揭示的是 extrapolation control 问题,而不是 floating-point overflow。
10. 与 P31 的关系
P32 参考 P31 的 paper-delivery spine:正文、完整 appendix、证据门禁、中文导读、PDF、clean arXiv source 和清晰结构图同步推进。算法上二者不同:P31 的 local 是冻结后的 top-\(k\) KRR inference;P32 v0.3 的 local 是 latent-coordinate law 所诱导的 affine prediction view。
主稿与 clean tarball 已独立编译,数学 convention、720-row evidence audit 与页面口径检查已通过,因此当前阅读包进入 public owner-review preview。这个 web preview 不自动触发 arXiv upload、venue submission 或正式 scientific release;这些仍需要 owner 的单独明确批准。