fix: step2 prompt 增加功能完整性要求 - Closes #75

新增规则 #9：要求 LLM 覆盖上下文包中的每个表格行和每条文字描述，确保不遗漏任何数据来源。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Merge pull request 'fix: [bug] Layer C QE Audit 持续 REJECT — 1/5 adequate 需提升至 ≥70% - 来自 #18 - Closes #75 ' (#77 ) from dev/issue-75-round2-ensemble-temp into main
2026-06-02 19:24:37 +08:00 · 2026-06-02 18:55:49 +08:00 · 2026-06-02 18:55:16 +08:00 · 2026-06-02 18:35:44 +08:00 · 2026-06-02 18:35:06 +08:00 · 2026-06-02 18:06:04 +08:00
6 changed files with 71 additions and 9 deletions
@@ -86,7 +86,8 @@ COVERAGE_TARGET = float(os.environ.get("IR_COVERAGE_TARGET", "0.95"))
 ENSEMBLE_TEMPERATURES = [
    float(os.environ.get("IR_ENSEMBLE_T1", "0.0")),
    float(os.environ.get("IR_ENSEMBLE_T2", "0.3")),
-    float(os.environ.get("IR_ENSEMBLE_T3", "0.7")),
+    float(os.environ.get("IR_ENSEMBLE_T3", "0.5")),
+    float(os.environ.get("IR_ENSEMBLE_T4", "0.7")),
 ]


@@ -186,6 +186,8 @@

 8. **开关关闭状态**：开关关闭时所有限制失效，这也必须作为一条规则输出（path: ["...", "开关关闭", "无限制"]）。

+9. **功能完整性要求（重要）**：上下文包中的每个表格行、每条文字描述、每个逻辑树路径都必须被至少一条规则覆盖。仔细检查上下文包，确保不遗漏任何数据来源。如果上下文包中有表格，每条表格行至少生成一条对应规则。
+
 {format_feedback}

 ## 输出格式
@@ -880,9 +880,9 @@ def run_ensemble_semantic_index(doc: dict) -> dict:
            if v:
                print(f"    {k}: {len(v)} 个问题")

-        # Feedback retry: re-run with coverage feedback (up to 2 retries, quality-gated)
+        # Feedback retry: re-run with coverage feedback (up to 3 retries, quality-gated)
        retry_count = 0
-        while retry_count < 2:
+        while retry_count < 3:
            feedback = _build_coverage_feedback(gaps)
            if not feedback:
                break
@@ -906,13 +906,16 @@ def run_ensemble_semantic_index(doc: dict) -> dict:
                            if src.get("section"):
                                retry_sections.add(src["section"])
                    print(f"  重试新增 sections: {sorted(retry_sections)}", flush=True)
-                    # Quality gate: only include retry if it improves coverage
+                    # Quality gate: include retry if it adds new sections or doesn't regress coverage
                    trial_indices = semantic_indices + [retry_result]
                    trial_merged = ensemble_merge(trial_indices)
                    trial_passed, trial_gaps = _quick_validate(trial_merged, doc, all_paths)
                    trial_warnings = len(trial_gaps.get("coverage_warnings", []))
                    trial_missing = len(trial_gaps.get("missing_table_rows", []))
-                    if trial_warnings < pre_warnings or trial_missing < pre_missing_rows:
+                    improved = trial_warnings < pre_warnings or trial_missing < pre_missing_rows
+                    no_regression = trial_warnings <= pre_warnings and trial_missing <= pre_missing_rows
+                    has_new_sections = len(retry_sections) > 0
+                    if improved or (no_regression and has_new_sections):
                        semantic_indices.append(retry_result)
                        merged = trial_merged
                        passed, gaps = trial_passed, trial_gaps
@@ -174,11 +174,25 @@ def _normalize_rule(rule: dict) -> dict:
    sources = rule.get("sources", [])
    valid_types = {"table", "text", "logic_tree"}

+    def _clean_section(val):
+        """Normalize section value: list→first element, ensure string."""
+        if isinstance(val, list):
+            return str(val[0]).strip() if val else ""
+        if isinstance(val, str):
+            return val.strip()
+        return str(val).strip() if val else ""
+
+    # Normalize section fields that might be lists (LLM format instability)
+    for s in sources:
+        sec = s.get("section")
+        if sec is not None:
+            s["section"] = _clean_section(sec)
+
    # try to infer a default section from the rule path
    default_section = ""
    for s in sources:
        sec = s.get("section", "")
-        if sec and sec.strip():
+        if sec and isinstance(sec, str) and sec.strip():
            default_section = sec.strip()
            break
    if not default_section:
@@ -192,7 +206,12 @@ def _normalize_rule(rule: dict) -> dict:
            if stype and stype not in valid_types:
                src["type"] = "text"
                stype = "text"
-            if stype in ("table", "text"):
+            if stype == "table":
+                if not src.get("section"):
+                    src["section"] = default_section
+                if src.get("row") is None:
+                    src["row"] = 0
+            elif stype == "text":
                if not src.get("section"):
                    src["section"] = default_section
    else:
@@ -512,6 +512,18 @@ class TestNormalizeRule:
        normalized = _normalize_rule(rule)
        assert "section" not in normalized["sources"][0]

+
+    def test_normalize_table_source_null_row(self):
+        """Table source with null row gets row=0 (defensive)."""
+        rule = {
+            "trigger": {"conditions": [{"signal": "x", "operator": "==", "value": "1"}]},
+            "sources": [
+                {"type": "table", "section": "3.1 功能", "row": None},
+            ],
+        }
+        normalized = _normalize_rule(rule)
+        assert normalized["sources"][0]["row"] == 0
+
    def test_normalize_source_invalid_type(self):
        """Invalid source types (LLM hallucinations) are normalized to text."""
        rule = {
@@ -538,3 +550,28 @@ class TestNormalizeRule:
        assert len(normalized["sources"]) == 1
        assert normalized["sources"][0]["type"] == "text"
        assert normalized["sources"][0]["section"] == "3.1 策略"
+
+    def test_normalize_section_is_list(self):
+        """Section field that is a list (LLM format bug) is normalized to string."""
+        rule = {
+            "trigger": {"conditions": [{"signal": "x", "operator": "==", "value": "1"}]},
+            "sources": [
+                {"type": "table", "section": ["状态", "系统设置"], "row": 1},
+                {"type": "text", "section": ["后台限制"], "text_snippet": "x"},
+            ],
+        }
+        normalized = _normalize_rule(rule)
+        assert normalized["sources"][0]["section"] == "状态"
+        assert normalized["sources"][1]["section"] == "后台限制"
+
+    def test_normalize_section_is_empty_list(self):
+        """Empty list section falls back to rule path."""
+        rule = {
+            "trigger": {"conditions": [{"signal": "x", "operator": "==", "value": "1"}]},
+            "path": "4.2 关闭流程 > decision",
+            "sources": [
+                {"type": "table", "section": [], "row": 1},
+            ],
+        }
+        normalized = _normalize_rule(rule)
+        assert normalized["sources"][0]["section"] == "4.2 关闭流程"
@@ -83,8 +83,8 @@ def test_output_dir_structure():


 def test_ensemble_temperatures_count():
-    """Should have exactly 3 ensemble temperatures."""
-    assert len(config.ENSEMBLE_TEMPERATURES) == 3
+    """Should have exactly 4 ensemble temperatures."""
+    assert len(config.ENSEMBLE_TEMPERATURES) == 4


 def test_max_tokens_is_int():
Author	SHA1	Message	Date
pzhang_zywl	440cd5812b	fix: step2 prompt 增加功能完整性要求 - Closes #75 CI / test (pull_request) Successful in 7s Details 新增规则 #9：要求 LLM 覆盖上下文包中的每个表格行和每条文字描述，确保不遗漏任何数据来源。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 19:24:37 +08:00
pzhang_zywl	55dcfc1b3e	Merge pull request 'fix: [bug] Layer C QE Audit 持续 REJECT — 1/5 adequate 需提升至 ≥70% - 来自 #18 - Closes #75 ' (#77 ) from dev/issue-75-round2-ensemble-temp into main CI / test (push) Successful in 9s Details	2026-06-02 18:55:49 +08:00
pzhang_zywl	4a8032665f	fix: ensemble 温度从 3 个增至 4 个增加多样性 - Closes #75 CI / test (pull_request) Successful in 8s Details 新增 t=0.5 温度变体，提高 ensemble 多样性以捕获更多功能单元。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 18:55:16 +08:00
pzhang_zywl	6536c7fa9d	Merge pull request 'fix: [bug] Layer C QE Audit 持续 REJECT — 1/5 adequate 需提升至 ≥70% - 来自 #18 - Closes #75 ' (#76 ) from dev/issue-75-retry-3 into main CI / test (push) Successful in 10s Details	2026-06-02 18:35:44 +08:00
pzhang_zywl	2cd02453ec	fix: step1 覆盖反馈重试增至 3 次 + 放宽质量门控 - Closes #75 CI / test (pull_request) Successful in 8s Details - 重试次数 2→3，增加 LLM 补全机会 - 质量门控放宽：新增 sections 且无回归即采纳，不只严格要求覆盖率下降 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 18:35:06 +08:00
pzhang_zywl	140e49342c	Merge pull request 'fix: [bug] step3 未防御 table source null row + Layer C QE Audit 100% 不合格 - 来自 #18 e2e - Closes #73 ' (#74 ) from dev/issue-73-fix-null-row into main CI / test (push) Successful in 8s Details	2026-06-02 18:06:04 +08:00
pzhang_zywl	93bbfe6029	fix: step3 _normalize_rule 将 table source 的 null row 转为 0 - Closes #73 CI / test (pull_request) Successful in 8s Details LLM 输出 table source 时 row 字段可能为 null，导致 Layer A schema 失败。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 18:05:28 +08:00
pzhang_zywl	6b1424b1c4	Merge pull request 'fix: [bug] step2 IR extraction 生成 list 类型 section 字段导致 conftest 崩溃 - 来自 #64 修复 - Closes #69 ' (#72 ) from dev/issue-69-fix-list-section into main CI / test (push) Successful in 12s Details	2026-06-02 17:45:37 +08:00
pzhang_zywl	efb5ed481e	fix: step3 _normalize_rule 处理 section 为 list 的 LLM 格式问题 - Closes #69 CI / test (pull_request) Successful in 9s Details LLM 输出 section 字段有时为 list 而非 string，导致 .strip() 崩溃。添加 _clean_section() 将 list→首元素 string，空 list 回退到 rule path。 Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>	2026-06-02 17:44:56 +08:00
pzhang_zywl	e54a221f34	Merge pull request 'fix: [test] conftest ir_data fixture 防御 LLM 产出的 list-type section - Closes #70 ' (#71 ) from test/issue-70 into main CI / test (push) Successful in 8s Details	2026-06-02 17:38:31 +08:00