mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-02 01:15:18 +08:00
feat(skills): add 习题课(最近发展区) (#1382)
* feat(skills): add ZPD-based exercise lesson skill * fix(skills): localize exercise lesson title in all workbench locales --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
This commit is contained in:
@@ -0,0 +1,60 @@
|
||||
# 习题课(最近发展区):local validation summary
|
||||
|
||||
## Scope and environment
|
||||
|
||||
- Date: 2026-09-05. Skill version: `0.2.0-experimental`.
|
||||
- Model used for the local classroom workflow: `deepseek:deepseek-v4-flash`.
|
||||
- Classroom checks used an existing local OpenMAIC runtime at `89268e5`, which includes the thinking-context fix `d5fca17` and an unrelated Feynman documentation commit. Neither change is included in this skill PR.
|
||||
- The PR branch starts from upstream `1e10f60`. The automated checks below were rerun there; the complete model-driven classroom workflow was **not** rerun on that clean runtime baseline.
|
||||
- All learner responses were constructed by the tester. Generated pages were inspected and locally corrected before the checks passed. This is not evidence of first-pass generation reliability or learning effectiveness.
|
||||
- The Chinese display name was finalized as **习题课(最近发展区)** after the classroom run. Its unchanged invocation id and new display title are checked by the discovery regression test.
|
||||
|
||||
## Observed classroom behavior
|
||||
|
||||
The existing three-activity exercise cycle was extended to six teaching pages without recreating the original diagnostic, supported-practice, or independent-check activities:
|
||||
|
||||
| Function | Material | Observed result |
|
||||
| --- | --- | --- |
|
||||
| Goal and route selection | Explain the goal and how to choose pages | Corrected the initial linear route into explicit conditional choices; reopened and visually checked the saved page |
|
||||
| Independent diagnosis | `3x + 5 = 20`, free response with a reason | Previous submitted answer, score, and feedback remained available after insertion/reordering |
|
||||
| Targeted review | Contrast one-sided and two-sided operations on `2a + 6 = 22` | Error explicitly labeled as an illustrative example, not an observed student mistake; mathematics and page rendering checked |
|
||||
| Optional supported practice | `4x + 3 = 19`; separate worked example `2m + 5 = 11` | Expanded the example, completed the two steps, and obtained the successful substitution check `19 = 19` |
|
||||
| Fresh independent check | `5x + 4 = 34` and added `2y + 7 = 25` | Two free-response inputs; no solutions or step prompts visible before submission; separate feedback and analysis after submission |
|
||||
| Reflection and next step | Initial attempt / help used / new-task performance / next step | Visible summary and manual record prompts; corrected a page-number typo and rechecked the saved page |
|
||||
|
||||
Both routes were exercised: diagnosis → independent check (skipping review/help), and review → supported practice. Navigation is learner/teacher-selected, not an automatic branching algorithm. Close the page-directory thumbnails during independent work so review material is not displayed beside the questions.
|
||||
|
||||
For the two-question submission, the simulated responses included `x = 6` and `y = 9`, with transformation reasons and substitution checks. The first explicitly stated that its solution had already been seen. The report returned 20/20 and separate comments; the first comment also said this was not independent completion and could not establish mastery. After closing the tab completely and reopening the classroom, both responses, scores, comments, and analyses remained available.
|
||||
|
||||
## Checks rerun on the isolated PR branch
|
||||
|
||||
- Skill frontmatter/name validation: passed.
|
||||
- `pnpm exec vitest run tests/agent-runtime/skills.test.ts tests/agent-runtime/skill-preload.test.ts tests/agent-runtime/skills-route.test.ts tests/agent-runtime/zpd-skill-discovery.test.ts`: **133 tests passed in 4 files**.
|
||||
- `pnpm check`: passed. The repository's normal Prettier configuration excludes Markdown; this is not a Markdown-content validation claim.
|
||||
- `pnpm lint`: passed with 0 errors and 18 warnings in unchanged upstream files; the new discovery test passed its focused lint check without warnings.
|
||||
- `pnpm exec tsc --noEmit --incremental false`: passed after linking the existing, lockfile-matched workspace dependency installation into the isolated worktree.
|
||||
- Node engine contract and i18n key alignment checks: passed.
|
||||
|
||||
These checks verify loading/integration and repository consistency, not instructional quality or learner-state persistence by themselves. The new test checks the actual discovered id, Chinese title, built-in source, and presence of both supporting references; it does not grade generated prose by keyword matching.
|
||||
|
||||
## Limitations and coverage boundaries
|
||||
|
||||
- The generated help page explicitly states that its temporary inputs do not survive a full document reload. It is not a cross-page learner record store; the lesson asks learners to retain their reflection separately.
|
||||
- Adding a question to an already-completed quiz temporarily displayed the old score against the new maximum (10/20). Starting a new attempt produced the two-question report above. Do not interpret the intermediate percentage as a decline in learning.
|
||||
- Non-mathematical transfer (claim–evidence–reasoning writing) received an independent instruction-level forward check, not a second model-generated classroom run. Revised dialogue rules also need broader future coverage; earlier dialogue tests are not relabeled as final-version passes.
|
||||
- The skill does not measure a stable ZPD score, establish durable mastery, or turn static optional help into automatic adaptive instruction.
|
||||
- TTS is explicitly outside this validation scope, not an outstanding acceptance gate.
|
||||
- No real-student learning outcome study was performed. Local session links, machine paths, credentials, and private learner data are not included in this public summary.
|
||||
|
||||
## CI follow-up: workbench names (2026-09-08)
|
||||
|
||||
The initial [CI job](https://github.com/THU-MAIC/OpenMAIC/actions/runs/33969660170/job/101315805424) failed one root unit test: `workbench-i18n.test.ts` reported that `zone-of-proximal-development` had no `zh-CN` display title. Skill frontmatter discovery had passed, but the built-in menu uses a separate workbench translation registry. The earlier four-file check did not cover that registry, and the general i18n key-alignment check did not detect a key absent from every locale.
|
||||
|
||||
The repair adds the title to both base languages and all ten locale overlays. It preserves the chosen Simplified Chinese name, the skill id, and the existing translation fallback behavior. The skill-specific regression now checks each locale's own copy and the actual menu-title resolver, so a missing overlay cannot pass by falling back to English or Simplified Chinese.
|
||||
|
||||
- Reproduced the original failing test locally before the change; it passed after the repair.
|
||||
- Expanded focused check, including `workbench-i18n.test.ts`: **236 tests passed in 5 files**.
|
||||
- Repository formatting, focused lint, TypeScript, and general i18n alignment checks passed.
|
||||
- The first default-concurrency full local run had one failure in the unchanged material-search character-budget test; that test file passed all 31 tests when run alone. A subsequent full run with `--maxWorkers=4` passed. No search implementation, time budget, test exclusions, or CI workflow was changed to obtain this result.
|
||||
|
||||
The original classroom evidence above remains dated to its actual runtime and has not been relabeled as a new classroom run.
|
||||
@@ -0,0 +1,67 @@
|
||||
# 习题课(最近发展区):本地产品检查
|
||||
|
||||
在真实 OpenMAIC 工作台选择“习题课(最近发展区)” Skill(内部标识 `zone-of-proximal-development`),使用现有配置的模型测试。测试输入均由测试者构造,不是学生学习数据。记录 Skill 版本、模型和运行基线;每项报告实际结果,不按关键词计分。会话和课堂链接可保留在私有测试记录,不作为公开复现的前提。
|
||||
|
||||
已有检查的范围、结果和限制见 [本地实测摘要](LOCAL_VALIDATION.md)。
|
||||
|
||||
## 名称与加载集成检查
|
||||
|
||||
内置 Skill 的菜单名称来自工作台多语言文案,不只来自 `SKILL.md` 的 `metadata.title`。新增或改名时同时核对 `lib/i18n/workbench.ts` 和各语言的 `workbench-locales` 文案,并运行以下检查;通用语言键对齐检查不能替代这项检查。
|
||||
|
||||
```bash
|
||||
pnpm exec vitest run tests/workbench/workbench-i18n.test.ts tests/agent-runtime/zpd-skill-discovery.test.ts tests/agent-runtime/skills.test.ts tests/agent-runtime/skill-preload.test.ts tests/agent-runtime/skills-route.test.ts
|
||||
```
|
||||
|
||||
除名称非空外,确认各语言有自己的文案,简体中文显示名保持“习题课(最近发展区)”。不要用修改翻译回退规则或跳过测试来掩盖缺失的名称。
|
||||
|
||||
## 对话:帮助后是否把任务交还学习者
|
||||
|
||||
新会话选择本 Skill,发送:
|
||||
|
||||
> 请在当前对话里带我练习一元一次方程,不创建课堂。我学过加减乘除。先给我一道短题,我自己试。
|
||||
|
||||
根据真实题目提交一个可解释的错解或卡点。随后明确说明缺少的知识,例如“我不知道为什么等号两边要做同样的运算”。观察系统是否补充必要教学、保留一道由测试者作答的任务并等待。
|
||||
|
||||
自己作答后,观察系统是否减少帮助并给出没有提前暴露答案的新题。再答新题,检查记录是否准确区分独立、受助与新任务表现。为保证真实性,下一轮输入随模型实际题目变化,不预先伪造已发生对话。
|
||||
|
||||
## 对话:已经会做与支持仍不足
|
||||
|
||||
用两个新会话分别检查:
|
||||
|
||||
> 我能独立解 3x+5=20:两边减5得3x=15,再除以3得x=5,代回左边等于20。请用最近发展区的方式给我下一步挑战,不创建课堂。
|
||||
|
||||
> 我想学一元一次方程,但还分不清乘法和加法。请用最近发展区的方式在对话里带我开始,不创建课堂。
|
||||
|
||||
前者应提供适度新挑战或一个必要澄清,避免重复完整入门示范。后者应缩小到合适的先备任务,不能只升级方程提示。两者均不能以一句回应计算稳定能力或发展区分数。
|
||||
|
||||
检查一轮中的多个问题是否彼此泄露答案。例如先问“2x是加还是乘”,随后又说“先乘2再加3”,不能将首问回答记为未受提示的证据。
|
||||
|
||||
## 真实练习循环
|
||||
|
||||
选择本 Skill,提出三页、中文、约十分钟的一元一次方程测试课:独立尝试、可自选帮助的练习、新题独立作答。使用真实页面生成与课堂预览。本轮教学流程测试不包含TTS验证,也不据此宣称音频通过验收。
|
||||
|
||||
检查 Skill 实际加载,三页实际保存;首题与新题在提交前不暴露答案;提示展开后仍有学习者动作;可以实际提交或操作;刷新后记录保留情况与页面声明一致。自选提示不等于自动诊断与撤架。任何失败均记录具体按钮、输入、操作和结果。
|
||||
|
||||
再用一份明确错解检查反馈,随后查看解答并重做,检查是否能修订和取回记录。重做成功只验证功能,不算首次独立表现。独立页、受助页与示范不要碰巧使用相同答案。
|
||||
|
||||
## 完整测验习题课
|
||||
|
||||
选择本 Skill,发送:
|
||||
|
||||
> 为学过正整数系数一元一次方程的初一学生整理一节20分钟测验习题课,要有独立诊断、必要讲评、分层练习、独立复核和收束,不只是题目集合。已会的学生能跳过带练,仍卡住的学生能回到适当帮助。请实际生成或完善课堂。
|
||||
|
||||
用真实页面清单核对目标、题组和讲评的关系;确认预设错例未冒充学生错误,初答与复核未被提前泄露,路线指引对应现有页面。实际走一次需要帮助的路线和一次跳过带练的路线,提交一道新增复核题并检查取回。不能把生成完成、静态自选路线、自动分流和学习效果混成一个“通过”。
|
||||
|
||||
再用非数学任务前向检查,例如:
|
||||
|
||||
> 把学生已学过的观点—证据—理由写作,整理成一节20分钟测验习题课。材料只有他们以前写过的短段落,不要替他们写答案。
|
||||
|
||||
检查是否能使用真实作品作初始证据、组织针对性讲评,让学生自己修订,并用新段落或新任务复核;不能把方程的步骤机械复制到写作。
|
||||
|
||||
## 理论主题边界
|
||||
|
||||
不选择 Skill,在新会话发送:
|
||||
|
||||
> 用150字解释维果茨基的最近发展区是什么,不要创建课堂,也不要安排练习。
|
||||
|
||||
应回答理论本身;不据这项测试认定其他输出的理论归因或教学效果已经验证。
|
||||
Reference in New Issue
Block a user