Compare commits

...
7 Commits
Author SHA1 Message Date
gurix 1474b0c66e Add goal-state skill: persistent, evidence-based goal tracking
File-based goal tracking (goal-state.yaml) for long or multi-session
work. Ordered (Sisyphus) and flexible modes, evidence-gated completion,
headless-safe (pi-web): no dialogs, no turn tricks.
2026-10-04 21:38:41 +02:00
Markus Graf 2bce5799dc Replace honest-reviewer with code-review skill
- Remove honest-reviewer (Socratic/devil's-advocate approach)
- Add code-review skill: high-precision review based on current best
  practices (confidence scoring >= 80, validate-before-report, context
  gathering beyond the diff, strict output contract, GLM/open-weight
  guardrails)
- dev-workflow: point review step at code-review, findings instead of
  questions; also adopt issue-based plan storage with explicit user
  approval gate
2026-09-28 21:42:01 +02:00
Markus Graf d22f69f47e Fix zai-usage.sh: export KEY so the Python fetch sends the auth header 2026-09-14 07:25:58 +02:00
gurix d155c0247e Add pi-zai-usage skill: query Z.ai GLM Coding Plan usage 2026-09-14 07:18:55 +02:00
Markus Graf 109e09cbf1 Merge remote origin/main into local main 2026-09-07 13:26:18 +02:00
Markus Graf a9be4bd4ae Add personal skills collection: dev-workflow and honest-reviewer 2026-09-07 12:15:34 +02:00
Markus Graf 593bbf20e9 chore: initial commit with honest-reviewer skill 2026-09-07 11:37:52 +02:00
8 changed files with 331 additions and 1 deletions
+6 -1
View File
@@ -1,2 +1,7 @@
# skills # Personal Skills Collection
This repository contains my personal collection of custom skills for the Pi coding agent.
## Author
Markus Graf <info@markusgraf.ch>
+9
View File
@@ -0,0 +1,9 @@
# Developer Skills Collection
This directory contains my personal collection of skills for software development.
For general development practices and minimalist coding, see [ponytail](https://github.com/DietrichGebert/ponytail).
For GitLab operations, use the [GitLab skills](https://github.com/gaodes/pi-gitlab/tree/main/skills) from @gaodes/pi-gitlab/skills/.
For persistent, evidence-based goal tracking across sessions, see [goal-state](goal-state/SKILL.md).
+121
View File
@@ -0,0 +1,121 @@
---
name: code-review
description: Reviews a branch, diff, or merge request like a skeptical senior engineer, reporting only high-confidence findings. Gathers context beyond the diff (intent, callers, project guidelines), validates every finding before reporting, and scores confidence (reports only >= 80/100). Covers correctness, contracts, security, and test coverage. Use before merging — or when the user says "review", "code review", "review my branch", or during the dev-workflow review step.
---
# Code Review
You review code changes like a senior engineer whose only currency is trust: one
false positive costs more than ten missed nits. In the wild, developers reject
over half of all automated review comments — your job is to be the exception.
## Prime directive
Report only findings you have **validated in the code**. If you cannot name the
concrete input, state, or code path that triggers a problem, do not report it.
"No issues found" is a valid, respected outcome — never invent findings to seem
thorough, never pad with style opinions.
This skill is designed to run with **fresh context** (e.g. in a subagent). If the
current session authored the changes it is about to review, say so and recommend
a fresh-context review instead: the assumptions made while writing code carry
over into its review.
## Workflow
### Phase 1 — Scope
1. Identify the base branch (`main`/`master`/`develop`). Range:
`git diff <base>...HEAD`, `git log <base>..HEAD`.
For an MR/PR on request: fetch the diff (`glab mr view <id> --raw`,
`gh pr diff <n>`) and apply the same workflow.
2. If there are no changes to review, say so and stop.
3. If the diff is large (roughly >600 changed lines): review in passes, one
group of related files at a time, prioritizing source code over tests, docs,
and generated files. State explicitly which files you did not fully review.
### Phase 2 — Context (the diff alone lies)
1. Read the **intent**: branch name, commit messages, MR/issue description if
available. State the intent in one sentence before judging the code.
2. Read project guidelines: `AGENTS.md`, `AGENT.md`, `CLAUDE.md`,
`CONTRIBUTING.md` at the repo root and in changed directories. Only enforce
rules that are actually written there — and quote the exact rule.
3. Trace the change into the codebase: open the full files around the hunks,
grep callers of changed functions, follow the data flow end to end. Do not
judge code you have not seen in context.
### Phase 3 — Review passes
Run the passes in this order; spend effort where the real defects live:
1. **Correctness and logic** — wrong conditions, off-by-one, inverted logic,
unhandled error paths (null, empty, timeout, retry), resource leaks, race
conditions, wrong state transitions.
2. **Contract breaks** — changed signatures or behavior vs. callers; API, DB,
or schema mismatches; serialization changes.
3. **Security at trust boundaries** — user input, files, network, env vars,
secrets in logs or history, auth checks, injection, unsafe deserialization.
4. **Tests** — which changed behavior has no test? Do new tests assert
outcomes, or merely execute code?
5. **Cheap checks** — if quick, run tests, linter, typecheck on the branch.
Report pass/fail in one line.
### Phase 4 — Validate and score
For every candidate finding, before it may be reported:
1. **Validate**: re-read the actual code and confirm the trigger path exists.
If the issue is handled elsewhere, drop the finding.
2. **Score confidence** 0–100: 0 = false positive · 50 = real but minor ·
75 = real and important · 100 = certain.
3. **Report only findings >= 80.** Exception: potential data loss or security
impact with lower confidence — report it with an explicit uncertainty note.
## Never report
- Pre-existing issues the branch does not make worse
- Code that only looks wrong but is actually correct
- Nitpicks a senior engineer would not flag
- Issues a linter or formatter will catch
- Style or quality subjectives; hypothetical "might be a problem" without a trigger
- Issues in code the diff does not touch
- Anything already silenced in code (lint-ignore comments)
## Output contract
Output exactly this structure — no preamble, no praise, no summary of what the
code does:
```
# Code Review: <branch>
Scope: <base>...HEAD, <n> files (+<a>/-<d>). Checks: <tests/lint/typecheck pass|fail|not run>.
## Findings
<in descending severity; omit the section entirely if there are none>
- **[blocker|should-fix]** `path:line` — <what is wrong>. Trigger: <concrete input/state/path>. Confidence: <n>/100.
<optional: one-sentence fix, only if obvious>
## Verdict
One line: safe to merge, or what must change first.
```
Severity definitions:
- **blocker** — data loss, security hole, certain crash, broken contract. Merge
must not happen.
- **should-fix** — real defect or unhandled failure mode that will bite.
There is no "nit" category — nitpicks are excluded entirely.
## Notes for open-weight models (GLM, Qwen, DeepSeek)
- Work through the phases strictly in order. Reason inside a phase (thinking
mode), but keep the final output to the contract above.
- Resist severity inflation: a finding without a concrete trigger scenario is
noise, not signal.
- Hard cap: at most 10 findings. If you found more, report the 10 highest-
severity ones and state how many you dropped.
- If context is missing to validate a candidate, drop the candidate — do not
report guesses.
+46
View File
@@ -0,0 +1,46 @@
---
name: dev-workflow
description: Personal development workflow for software projects.
---
# Dev Workflow
This skill enforces my personal development workflow for software projects.
## Branch Strategy
- **Main branch**: `master` or `main` (production)
- **Development**: Always on feature branches off main
- **New work**: Features, fixes, and chores each get their own branch
## Workflow
### 1. Research / Discussion / Plan
Discuss the task, idea, or bug with the coding agent. The outcome is an **implementation plan** that includes:
- Summary of research and discussion
- Implementation plan broken into discrete steps
Create a GitLab/Git issue for the plan. The issue title starts with the feature slug, and the description contains the full plan. **WAIT for explicit user approval before creating the issue.**
### 2. Implementation
**Only proceed if the user explicitly approves the plan** (e.g., `go`, `approved`, `start`). Do not start implementing on your own. Before starting, the agent reads `AGENT.md` (project guidelines). Implementation proceeds step-by-step according to the plan. Each step is committed individually.
### 3. Review
After implementation is complete, a subagent with **fresh context** reviews the code and implementation using `code-review`.
### 4. Iteration
The agent fixes any findings the reviewer reported. This cycle continues until the reviewer subagent approves.
### 5. Merge Request
Create an MR/PR and inform the user.
## Language Rule
- **Plans and artifacts**: Always in **English**
- **User discussions**: Any language
+76
View File
@@ -0,0 +1,76 @@
---
name: goal-state
description: Persistent, evidence-based goal tracking for long or multi-session work via a goal-state.yaml file in the project. Use when the user says "goal", "/goal", "setze ein Ziel", "track this as a goal", wants ordered step-by-step execution with done criteria (Sisyphus mode), or when a goal-state.yaml already exists. Works headless (pi-web) — file-based, no dialogs or TUI needed.
---
# Goal State
Track a goal in `goal-state.yaml` in the project root. The file is the single source of truth: it survives sessions, compaction, and crashes. Every claim of progress needs recorded evidence.
## The state file
```yaml
version: 1
status: active # active | paused | complete | cancelled
mode: ordered # ordered | flexible
objective: >
One-sentence outcome.
requirements: # the goal is only complete when ALL are met
- id: R1
text: Verifiable completion requirement.
met: false
evidence: "" # how it was verified (command, output, file path)
steps:
- id: S1
text: What to do.
done_when: How to verify this step is done.
status: pending # pending | in_progress | done | skipped
evidence: ""
next: S1 # ordered mode only: the one step to work on
log: # newest last, keep at most 20 entries
- "2026-01-15T10:00Z goal created"
```
- **ordered** (Sisyphus): steps are executed strictly in order, one at a time. `next` names the only step that may be worked on.
- **flexible**: steps may be done in any order; requirements still gate completion.
## Starting a goal
1. Discuss the objective. Derive **requirements** — each must be independently verifiable (test suite passes, file exists, report contains section X, command exits 0).
2. For ordered mode, break the work into numbered steps, each with a concrete `done_when`. For flexible mode, steps are optional — requirements alone are fine for small goals.
3. Show the proposed file to the user and **wait for confirmation**. Then write `goal-state.yaml` and log the creation.
Shortcut: if the user gives a complete objective with explicit steps, propose immediately without a long discussion.
## Working on a goal
1. Read `goal-state.yaml` before doing anything else.
2. **Ordered:** set the `next` step to `in_progress`, do it, verify its `done_when`. **Flexible:** pick any pending step.
3. Record evidence and set `status: done`, advance `next`, append a log line — **write the file immediately after each step** (crash safety).
4. Blocked? Set `status: paused` with the reason in the log and tell the user what is needed.
5. If the project uses the dev-workflow, each step gets its own commit, per that workflow.
## Evidence rules
- `met: true` / `status: done` **only with evidence recorded**: the command that was run and its result, a file path, or a test name. "I am confident" is not evidence.
- When unsure whether something still holds, re-run the verification instead of trusting the file.
- Evidence from earlier sessions counts only if the file records it.
## Resuming
- If a session starts (or the skill is loaded) and `goal-state.yaml` has `status: active`, give a two-line status (objective, next step) and continue with `next` — ask only if the situation is ambiguous.
- `goal status` → print a compact view from the file: status, requirements (met/unmet), steps (done/pending), next step. No narration beyond that.
## Completing, changing, cancelling
- **Complete:** only when every requirement has `met: true` with evidence. Set `status: complete`, log it, and report a summary of objective + evidence to the user. Suggest `mv goal-state.yaml goal-state-done-<date>.yaml` (or deletion).
- **Change:** adjust steps/requirements, append a log entry describing what changed and why.
- **Cancel:** set `status: cancelled`, keep the file unless the user wants it deleted.
## Rules
- One active goal per project.
- Never skip steps in ordered mode; never work ahead of `next`.
- Never claim completion without evidence in the file.
- Write the file after every mutation — the file, not memory, is the progress.
- Keep the file small: log capped at 20 entries, no prose beyond the objective.
+3
View File
@@ -0,0 +1,3 @@
# Integration Skills Collection
Skills that integrate external services and APIs with the Pi coding agent.
+16
View File
@@ -0,0 +1,16 @@
---
name: zai-usage
description: Query current Z.ai GLM Coding Plan quota and usage (5-hour/weekly credits, model tokens, MCP calls). Use when the user asks about their Z.ai usage, remaining credits, quota, or rate-limit status.
---
Run the script once and report the result:
```bash
bash ~/.pi/agent/skills/zai-usage/scripts/zai-usage.sh # last 24h window
bash ~/.pi/agent/skills/zai-usage/scripts/zai-usage.sh 168 # last 7 days
```
- Credentials come from `~/.pi/agent/auth.json` (`zai.key`) — never print the key.
- Endpoints: `api.z.ai/api/monitor/usage/{model-usage,tool-usage,quota/limit}` (same as Z.ai's Claude Code plugin).
- Present 5-hour and weekly credits (used/remaining/reset) plus model token totals in a small table.
- If the key is missing, tell the user to run `/login zai` in pi first.
+54
View File
@@ -0,0 +1,54 @@
#!/usr/bin/env bash
# Query Z.ai GLM Coding Plan usage (same endpoints as the Claude Code glm-plan-usage plugin).
# Usage: zai-usage.sh [hours] (default: 24h lookback window)
set -euo pipefail
HOURS="${1:-24}"
AUTH_JSON="$HOME/.pi/agent/auth.json"
export KEY="$(python3 -c "import json;print(json.load(open('$AUTH_JSON'))['zai']['key'])")"
BASE="https://api.z.ai/api/monitor/usage"
NOW="$(date -u +'%Y-%m-%d %H:%M:%S')"
START="$(date -u -d "$HOURS hours ago" +'%Y-%m-%d %H:%M:%S')"
Q="startTime=$(python3 -c "import urllib.parse,sys;print(urllib.parse.quote(sys.argv[1]))" "$START")&endTime=$(python3 -c "import urllib.parse,sys;print(urllib.parse.quote(sys.argv[1]))" "$NOW")"
fetch() { curl -sS -m 20 "$1" -H "Authorization: $KEY" -H 'Accept-Language: en-US'; }
python3 - "$BASE" "$Q" <<'EOF'
import json, subprocess, sys
from datetime import datetime, timezone
base, q = sys.argv[1], sys.argv[2]
def fetch(path, params=""):
out = subprocess.run(["curl", "-sS", "-m", "20", f"{base}/{path}{params}",
"-H", f"Authorization: {__import__('os').environ.get('KEY','')}"],
capture_output=True, text=True, env={**__import__('os').environ}).stdout
return json.loads(out).get("data")
# Quota / limits
data = fetch("quota/limit")
level = data.get("level", "?")
print(f"Plan: {level.upper()}\n")
for lim in data.get("limits", []):
if lim["type"] != "CREDIT_LIMIT":
continue
unit = "5-hour" if lim["unit"] == 3 else ("weekly" if lim["unit"] == 6 else f"unit{lim['unit']}")
reset = datetime.fromtimestamp(lim["nextResetTime"] / 1000, tz=timezone.utc)
delta = reset - datetime.now(timezone.utc)
hrs = int(delta.total_seconds() // 3600); mins = int(delta.total_seconds() % 3600 // 60)
when = f"in {hrs}h{mins:02d}m" if hrs else f"in {mins}m"
print(f"{unit:>7} credits: {lim['currentValue']:>6} / {lim['usage']} used ({lim['percentage']}%)"
f" — remaining {lim['remaining']}, resets {when}")
# Model usage in window
data = fetch("model-usage", f"?{q}")
tot = data.get("totalUsage", {})
print(f"\nLast window: {tot.get('totalModelCallCount', 0)} model calls")
for m in tot.get("modelSummaryList", []):
print(f" {m['modelName']}: {m['totalTokens']:,} tokens")
tools = fetch("tool-usage", f"?{q}").get("totalUsage", {})
ncalls = tools.get("totalNetworkSearchCount", 0) + tools.get("totalWebReadMcpCount", 0) + tools.get("totalZreadMcpCount", 0)
print(f"MCP tool calls: {ncalls}")
EOF