fix: ensemble 温度从 3 个增至 4 个增加多样性 - Closes #75

新增 t=0.5 温度变体，提高 ensemble 多样性以捕获更多功能单元。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Merge pull request 'fix: [bug] Layer C QE Audit 持续 REJECT — 1/5 adequate 需提升至 ≥70% - 来自 #18 - Closes #75 ' (#76 ) from dev/issue-75-retry-3 into main
2026-06-02 18:55:16 +08:00 · 2026-06-02 18:35:44 +08:00 · 2026-06-02 18:35:06 +08:00 · 2026-06-02 18:06:04 +08:00 · 2026-06-02 18:05:28 +08:00 · 2026-06-02 17:45:37 +08:00
5 changed files with 29 additions and 8 deletions
@@ -86,7 +86,8 @@ COVERAGE_TARGET = float(os.environ.get("IR_COVERAGE_TARGET", "0.95"))
 ENSEMBLE_TEMPERATURES = [
    float(os.environ.get("IR_ENSEMBLE_T1", "0.0")),
    float(os.environ.get("IR_ENSEMBLE_T2", "0.3")),
-    float(os.environ.get("IR_ENSEMBLE_T3", "0.7")),
+    float(os.environ.get("IR_ENSEMBLE_T3", "0.5")),
    float(os.environ.get("IR_ENSEMBLE_T4", "0.7")),
 ]
@@ -880,9 +880,9 @@ def run_ensemble_semantic_index(doc: dict) -> dict:
            if v:
                print(f"    {k}: {len(v)} 个问题")
-        # Feedback retry: re-run with coverage feedback (up to 2 retries, quality-gated)
+        # Feedback retry: re-run with coverage feedback (up to 3 retries, quality-gated)
        retry_count = 0
-        while retry_count < 2:
+        while retry_count < 3:
            feedback = _build_coverage_feedback(gaps)
            if not feedback:
                break
@@ -906,13 +906,16 @@ def run_ensemble_semantic_index(doc: dict) -> dict:
                            if src.get("section"):
                                retry_sections.add(src["section"])
                    print(f"  重试新增 sections: {sorted(retry_sections)}", flush=True)
-                    # Quality gate: only include retry if it improves coverage
+                    # Quality gate: include retry if it adds new sections or doesn't regress coverage
                    trial_indices = semantic_indices + [retry_result]
                    trial_merged = ensemble_merge(trial_indices)
                    trial_passed, trial_gaps = _quick_validate(trial_merged, doc, all_paths)
                    trial_warnings = len(trial_gaps.get("coverage_warnings", []))
                    trial_missing = len(trial_gaps.get("missing_table_rows", []))
-                    if trial_warnings < pre_warnings or trial_missing < pre_missing_rows:
+                    improved = trial_warnings < pre_warnings or trial_missing < pre_missing_rows
                    no_regression = trial_warnings <= pre_warnings and trial_missing <= pre_missing_rows
                    has_new_sections = len(retry_sections) > 0
                    if improved or (no_regression and has_new_sections):
                        semantic_indices.append(retry_result)
                        merged = trial_merged
                        passed, gaps = trial_passed, trial_gaps
@@ -206,7 +206,12 @@ def _normalize_rule(rule: dict) -> dict:
            if stype and stype not in valid_types:
                src["type"] = "text"
                stype = "text"
-            if stype in ("table", "text"):
+            if stype == "table":
                if not src.get("section"):
                    src["section"] = default_section
                if src.get("row") is None:
                    src["row"] = 0
            elif stype == "text":
                if not src.get("section"):
                    src["section"] = default_section
    else:
@@ -512,6 +512,18 @@ class TestNormalizeRule:
        normalized = _normalize_rule(rule)
        assert "section" not in normalized["sources"][0]
    def test_normalize_table_source_null_row(self):
        """Table source with null row gets row=0 (defensive)."""
        rule = {
            "trigger": {"conditions": [{"signal": "x", "operator": "==", "value": "1"}]},
            "sources": [
                {"type": "table", "section": "3.1 功能", "row": None},
            ],
        }
        normalized = _normalize_rule(rule)
        assert normalized["sources"][0]["row"] == 0
    def test_normalize_source_invalid_type(self):
        """Invalid source types (LLM hallucinations) are normalized to text."""
        rule = {
@@ -83,8 +83,8 @@ def test_output_dir_structure():
 def test_ensemble_temperatures_count():
-    """Should have exactly 3 ensemble temperatures."""
+    """Should have exactly 4 ensemble temperatures."""
-    assert len(config.ENSEMBLE_TEMPERATURES) == 3
+    assert len(config.ENSEMBLE_TEMPERATURES) == 4
 def test_max_tokens_is_int():
Author	SHA1	Message	Date
pzhang_zywl	4a8032665f	fix: ensemble 温度从 3 个增至 4 个增加多样性 - Closes #75 CI / test (pull_request) Successful in 8s Details 新增 t=0.5 温度变体，提高 ensemble 多样性以捕获更多功能单元。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 18:55:16 +08:00
pzhang_zywl	6536c7fa9d	Merge pull request 'fix: [bug] Layer C QE Audit 持续 REJECT — 1/5 adequate 需提升至 ≥70% - 来自 #18 - Closes #75 ' (#76 ) from dev/issue-75-retry-3 into main CI / test (push) Successful in 10s Details	2026-06-02 18:35:44 +08:00
pzhang_zywl	2cd02453ec	fix: step1 覆盖反馈重试增至 3 次 + 放宽质量门控 - Closes #75 CI / test (pull_request) Successful in 8s Details - 重试次数 2→3，增加 LLM 补全机会 - 质量门控放宽：新增 sections 且无回归即采纳，不只严格要求覆盖率下降 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 18:35:06 +08:00
pzhang_zywl	140e49342c	Merge pull request 'fix: [bug] step3 未防御 table source null row + Layer C QE Audit 100% 不合格 - 来自 #18 e2e - Closes #73 ' (#74 ) from dev/issue-73-fix-null-row into main CI / test (push) Successful in 8s Details	2026-06-02 18:06:04 +08:00
pzhang_zywl	93bbfe6029	fix: step3 _normalize_rule 将 table source 的 null row 转为 0 - Closes #73 CI / test (pull_request) Successful in 8s Details LLM 输出 table source 时 row 字段可能为 null，导致 Layer A schema 失败。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 18:05:28 +08:00
pzhang_zywl	6b1424b1c4	Merge pull request 'fix: [bug] step2 IR extraction 生成 list 类型 section 字段导致 conftest 崩溃 - 来自 #64 修复 - Closes #69 ' (#72 ) from dev/issue-69-fix-list-section into main CI / test (push) Successful in 12s Details	2026-06-02 17:45:37 +08:00