Same-unit counterfactual 需要什么额外 coupling?
下一门禁是 formal impossibility result 与 non-vacuous partial-identification construction。
这个页面是 10 Paper Portfolio 的研究问题入口。研究问题本身是第一产物, 能回答多少不是第一优先级。这里不是论文标题清单,也不是 44 篇文章的承诺表; 问题可以先保持开放,再逐渐长出 evidence、SOTA pressure、objection、实验或论文。
当前不是平均推进 44 个槽位。下面先列出 WIP=4 的既有授权集合,再列出 P33、P34 两个不占 WIP 的正式 slot; 当前唯一下一科学门禁仍是 USL-03,USL-04 在其结果记录后进入。48h path 不要求给出答案; 它可以只是把问题为何值得问、最强异议和下一条反馈或证据路径说得更清楚。
下一门禁是 formal impossibility result 与 non-vacuous partial-identification construction。
下一门禁是冻结 pre-treatment estimand、cross-fitting protocol 与 matched experiment。
corrected-v4 已完成 integrity、objective 与 claim audit;当前只保留 clean k=100 的 scope-limited local contrast。v3 是 objective-unvalidated provenance,mechanism、Cauchy-specific、scale 与 general-baseline claims 均未建立。
v0.5 已对齐 USL01 response interface,UnitRec-Synth v0.1 的九项 controlled Gaussian whole-user checks 全部通过;real-world selector 仍未验证。冻结 v0.3 是 collapsed proxy diagnostic,其 historical HTML projection noindex 且 stale against v0.5,linked artifacts 仍公开。
Stage 1 四机制五 seed matrix 已完成 340/340,并分别通过 mixture 与 tail gate;五数据集 real Stage 1 完成 375/375,但不建立真实 conditional multimodality。Mandatory scale v0.2 完成 190/190,仍有 revision/full-host provenance hold;Stage 2 surface、external baselines 与独立 host reproduction 待完成。
Ordered C=1000 synthetic gate 显示 structured-label signal,random permutation 几乎抹除该信号;C=5000 slice 未保持 accuracy,local CPU exact scoring 慢于 softmax,且没有独立 Student-t tail gain。下一门禁是 C={100,1000,5000,10000} 三 seed、semantic/learned-code controls、强 label-embedding/MHE opponents,以及 exact-versus-fast Top-r decoding。P34 不占当前 WIP,也不改变 USL-03 → USL-04 顺序。
新问题不需要先答完下面五项才有资格进入。先把核心疑问留下来;这些提示只帮助它逐渐 变清楚。答不上来,可以继续保持 open,也不要为了进入论文流程而制造答案。
`RQxx` 是稳定问题槽位;`Pxx` / old `A2` / `B5` 只是候选或论文 handle。 A/B/C package role 必须等 evidence 成熟后再浮现。
| Slot | Research question | Handle | State | 48h evidence | Kill / merge rule |
|---|---|---|---|---|---|
| RQ01 | When should causal modeling start from observational data, and how does the answer change if heterogeneity is primitive? | P01 observational survey | review wedge | Outline plus source matrix with 30 core references. | Kill expansion if it becomes a generic causal inference method list. |
| RQ02 | Can structural causal mechanism interfaces improve LLM / agent counterfactual reasoning beyond prompt-only baselines? | old P01 / A1 | live | One benchmark task where mechanism structure changes answer or failure attribution. | Park if novelty is only prompt engineering around existing benchmarks. |
| RQ03 | Can agent causal reasoning be evaluated by decision deltas in real workflows? | P02 / old A2 | sprint | Independent recoding or broader sample for the 10-event decision-delta table. | Downgrade to internal eval if deltas are post-hoc or founder-specific. |
| RQ04 | Can collaboration traces become causal data for improving multi-agent workflows? | P03 / old A3 | live | Redaction-safe event schema over 20 events and one intervention contrast. | Park if privacy redaction destroys signal or examples stay anecdotal. |
| RQ05 | What failure-class semantics do agent runtime health probes need to avoid wrong escalation? | P04 / old B1 | sprint | One less self-referential replay case plus verified opponent metadata. | Downgrade to runbook/blog if taxonomy is standard SRE with no agent-runtime contribution. |
| RQ06 | Can a tabular foundation model operate on ultra-high-dimensional, variable-length sparse feature sets without a fixed maximum column axis, with recommendation as the first stress test? | P05 | live owner-review seed | Matched 10^4/10^5/10^6 sparse episodes with variable active counts, unseen IDs, schema drift, memory and latency telemetry. | Kill if ordinary sparse embeddings or a fixed vocabulary explain the result. |
| RQ07 | Can sparse human taste be converted into durable agent work changes? | P06 / old B3 | parked | Five real feedback events with artifact_delta or decision_delta tracking. | Park until real feedback changes downstream work. |
| RQ08 | Can app-less agent harnesses optimize communication bandwidth better than dashboard-first workflows? | P07 / old B4 | parked | Observable bandwidth rubric over three workflows. | Park if the metric remains taste-only or too subjective. |
| RQ09 | What state machine helps multi-agent paper production decide sprint / park / merge / kill? | P08 / old B5 | sprint | Convert top candidates into structured entries and prove work allocation changed. | Kill as paper if it adds bureaucracy without changing decisions. |
| RQ10 | Can Causality Primer become an agent skill curriculum with task-level evaluation? | P09 / old C1 | parked | One chapter becomes one skill exercise plus eval. | Keep as skill asset unless it gains research contribution. |
| RQ11 | Can LLM-derived semantic feature tokens and evidence-abducted Unit tokens support RecSys-style masked prediction across heterogeneous table schemas with exact column-order invariance? | P11 tabular foundation models | v1 architecture seed | Freeze semantic-column TFM overlap, cross-table parameter sharing, LLM training and adaptation contracts before a bounded multi-schema test. | Fold back into the review asset if this is an existing semantic-column TFM under Unit vocabulary or creates no discriminating result. |
| RQ12 | What agent profile state is minimally necessary for team intelligence? | old C3 / component of RQ04 | component | Before/after routing-error comparison. | Keep as Team OS background unless it creates independent evidence. |
| RQ13 | Does a kill matrix prevent AI-generated paper-shaped artifacts? | P12 / old C4 | component | Apply reviewer objections to 15 candidates and log changed decisions. | Keep merged into RQ09 unless it becomes independently measurable. |
| RQ14 | How can team agent logs preserve causal signal without leaking private context? | P13 / old C5 | parked | Redaction levels over three concrete log examples. | Park until publishable examples exist. |
| RQ15 | Can probability distributions be directly aggregated by a defensible information aggregation operator? | WEDGE-001 | live | Two-page SOTA/opponent memo with 10 closest related works. | Park if the operator reduces to known pooling/fusion or lacks advantage. |
| RQ16 | How can the first bounded DiscoSCM slice learn a token-modulated outcome mechanism from one factual observation per token? | Learning DiscoSCM · P16 supporting seed | flagship science | V2 logic contract and factual-law/non-identifiability/Cauchy/off-diagonal propositions; next run is a truth-isolated alternative-query benchmark. | Not an active submission package; require comparator-aware evidence and do not infer Layer 3 identification from factual-law recovery. |
| RQ17 | What is the minimal benchmark for DiscoSCM-style counterfactual / intervention reasoning? | DiscoSCM eval | open | Three tasks, baseline, scoring rule, failure taxonomy. | Park if tasks duplicate existing causal LLM benchmarks. |
| RQ18 | Can CausalQwen / CausalLLM evidence be rebuilt into a credible causal-reasoning package? | CausalLLM | open | Locate current evidence, dataset, baseline, and reproduction path. | Park if evidence cannot be reproduced. |
| RQ19 | When are population-level cycles projections of unit-level DAG mechanisms? | Primer Ch03 | open | Formal note plus examples separating projection cycles from equilibrium semantics. | Merge into Primer if no paper-level opponent appears. |
| RQ20 | Can robust prediction under distribution shift be reframed as mechanism preservation across heterogeneous units? | causal representation | open | Three SOTA opponents plus one toy shift example. | Park if it is a generic distribution-shift survey. |
| RQ21 | Can causal mechanism diagnosis improve LLM / agent failure attribution? | Causal intelligence | open | Two failure cases where mechanism diagnosis changes repair path. | Merge into RQ02/RQ03 if not independently sharp. |
| RQ22 | How should no-data causal reasoning be represented before observational evidence enters? | P01 source bucket | open | Source matrix examples from action, physics, social mechanism, and counterfactuals. | Merge into RQ01 unless it demands standalone treatment. |
| RQ23 | What changes when unit-specific treatment-response mechanisms become primary objects beyond HTE / ITE? | P23 / HCGM causal-effect | preflight | Upstream theory, robustness, IHDP and WAWS evidence plus HTE/ITE comparison. | Reposition if the object-level distinction adds no value beyond direct robust or representation learners. |
| RQ24 | Can PyWhy / DoWhy-style workflows become executable teaching/evaluation assets for causal intelligence? | Primer tooling | open | One model-identify-estimate-refute case connected to Primer. | Keep as course/tooling asset unless evaluation becomes novel. |
| RQ25 | Can old manuscript assets be upgraded through a standard SOTA/opponent pass? | old draft protocol | open | Inventory five old drafts and run one complete upgrade/park memo. | Merge into RQ09 if it remains production protocol. |
| RQ26 | Which research problems should gong personally deep-touch, given a 2-3 slot attention budget? | owner attention | open | Taste/novelty/owner-need rubric applied to all live slots. | Keep internal unless it becomes RQ09 evidence. |
| RQ27 | Can a source matrix prevent survey papers from becoming literature lists? | P01 process | open | Build source matrix for RQ01 and record how it changes outline. | Merge into RQ09 if it is only workflow evidence. |
| RQ28 | Can a single-paper entry site improve public feedback and reviewer readiness? | public projection | open | One P01 or RQ09 entry updated from placeholder to real public-safe page. | Keep as delivery asset unless feedback effect is measured. |
| RQ29 | What is the publication boundary between WeHub internal evidence and public research claims? | research integrity | open | Boundary memo using P02/P04/RQ14 as cases. | Merge into RQ14 if it becomes only log privacy. |
| RQ30 | How should WeHub mine new research questions continuously instead of waiting for brainstorming bursts? | question engine | open | Daily question intake protocol plus one week of changes to board state. | Keep as internal loop unless it produces publishable method evidence. |
| RQ31 | Does a mathematically valid learnable location–scale Cauchy–KL kernel provide value beyond matched RBF and fixed-scale Cauchy controls? | P31 / learnable Cauchy–KL kernel | corrected-v4 complete + claim audited | Integrity and objective repair pass for this artifact; the audit retains only scope-limited clean k=100 local support (27/27 lower means, 26/27 exploratory intervals excluding zero, 131/135 seed wins). No mechanism, Cauchy-specific, input-dependent-scale, or general-baseline claim is established; deep-vs-fixed remains grid_sensitivity_hold and Candidate U remains prospective. | Run fixed-grid sensitivity or objective-valid mechanism experiments only for a claim actually pursued; otherwise retain this scoped result and preserve corrected-v3 as objective-unvalidated provenance. |
| RQ32 | Can an evidence-conditioned stable latent-coordinate law turn globally nonlinear factual prediction into coordinate-conditioned local affine views with an exactly marginalized stable predictive law, and what scientific type should that coordinate have? | P32 / UALBM | v0.5 controlled unit semantics validated / real-world unvalidated | The realized identity-bearing unit now directly selects the response-law member, and UnitRec-Synth v0.1 passes all nine controlled Gaussian checks. Frozen v0.3 tabular runs remain collapsed proxy diagnostics; their historical HTML owner-review projection is noindex and stale against v0.5, while linked PDF/tar artifacts remain public until server policy changes. | Pressure-test closest opponents and a declared-unit real-world selector; keep grouped/deduplicated and robustness gates scoped to the v0.3 proxy, and downgrade to a technical note if only standard ingredients survive review. |
| RQ33 | Can a shared-encoder Student-t mixture density network improve conditional-density quality on tabular regression by separating heavy-tail robustness from input-dependent multimodality, with Gaussian and Cauchy as matched special cases? | P33 / Shared-Encoder Student-t MDN | formal slot / Stage 1 mechanism gate passed / Stage 2 pending | The four-mechanism synthetic matrix completed 340/340 runs and separates mixture and tail effects; the five-dataset real panel completed 375/375 runs but supports no real multimodality claim. Scale v0.2 completed 190/190 with a revision/full-host provenance hold. | Run the Stage 2 phase diagram, categorical Ames/CCPP, external baselines, Faithful/beta-NLL controls and independent-host reproduction; retain the conditional benchmark fallback When Do Student-t Mixture Density Networks Help?. |
| RQ34 | Can fixed-complexity density coding reduce learned output-head parameters for extreme multiclass prediction when labels have a defensible order? | P34 / MIXCODE | formal slot / ordered-label signal / scaling pending | The bounded synthetic gate shows an ordered-label signal and a sharp random-permutation failure; C-scaling, exact-ranking latency, Student-t-specific gain, and real-data evidence remain unresolved. | Require three seeds, semantic/learned code controls, strong embedding/MHE opponents, and exact-versus-fast Top-r decoding; retain only a structured-label benchmark if the general claim fails. |
Unit Selection Learning 的十个 registered AAAI-27 课题拥有独立身份, 与上述 34 个 RQ 槽位并行。它们共享同一 unit-belief 理论边界,但每篇有独立可证伪 claim。 进入 USL 十篇完整入口 →