The job of a software engineer is to automate things. Not just to type code. Not just to close tickets. The job is to notice a repeated operation, understand the judgment inside it, and turn that judgment into a system that runs with less human attention next time.

AI agents change the unit of automation. Before, we mostly automated deterministic work: scripts, tests, deployments, migrations, formatters, dashboards. Now we can also automate parts of the engineering loop itself: planning, exploration, implementation, evidence gathering, and reporting. But only if the work is specified in a way an agent can execute and in a way we can evaluate.

That is why goal writing matters. A goal is not a motivational sentence. It is an interface between human intent and autonomous work. A good goal encodes the objective, constraints, evidence standard, risk model, time budget, and review surface. A bad goal leaks all of that back into chat, where the human has to babysit the run.

task promptwhat the user askedgoal writerskill under testGOAL.mdcontract for workexecutor45 min in repojudgescores artifactsreferencehidden lessonsrepeat across skills, models, harnessesreview findings become better skills

Automate the job

A lot of engineering automation fails because it automates the visible keystrokes instead of the actual job. The visible keystroke is “edit this file.” The job is usually something bigger: preserve behavior while optimizing, reproduce a UI without faking it, migrate a system without data loss, diagnose an intermittent bug, or explore three approaches and report which one survived contact with reality.

Human engineers carry hidden policies for this. We know when to stop and benchmark. We know when screenshots are not evidence. We know when a passing unit test is too narrow. We know when a clever refactor is risky because the original code had weird edge cases. These policies are the real thing worth automating.

Agentic software engineering therefore needs a second artifact next to code: an explicit operating contract for the run. That contract should tell the agent not only what to build, but how to decide, what to preserve, what evidence to collect, and what to show the human at the end.

Goals are interfaces

Think of a goal like an API. A vague goal is an untyped function with side effects. It accepts anything, returns whatever happened, and forces the caller to inspect logs to understand whether it worked. A strong goal has a schema.

Goal quality
PartWeak goalStrong goal
Objectivemake it betteroptimize scanner without changing results
Scopework on the apptouch only scanner + benchmark harness unless needed
Evidenceseems fasterbefore/after timings plus correctness comparison
Riskimplicitlist correctness risks before changing algorithm
Outputsummary in chatreport file, commands run, unresolved items
Fallbacknoneif blocked, preserve evidence and next best actions

This framing makes goal writing feel less like prompt craft and more like interface design. You are defining the boundary conditions for an autonomous worker. The goal should be precise enough that the worker can make local decisions, but not so prescriptive that it prevents useful exploration.

goal = { objective: concrete_outcome, context: files + constraints + user_taste, budgets: time + tokens + risk, evidence: tests + benchmarks + screenshots + reports, autonomy: what to decide without asking, stop_rules: when to report instead of thrash, review_surface: artifacts the human can inspect }

The skill

A goal-writing skill is a reusable procedure for turning a rough user request into that contract. It is not specific to one repository. It is a general compiler from intent to execution plan.

The skill should ask a consistent set of questions internally:

  • What is the real deliverable? Code, report, benchmark, diagnosis, prototype, migration, reproduction, or decision memo?
  • What must not be faked? No screenshot-only UI, no invented benchmark, no speculative behavior, no placeholder passing as implementation.
  • What evidence proves progress? Tests, source links, before/after metrics, manual QA, logs, saved artifacts, or comparison against reference behavior.
  • Where should autonomy be spent? Safe exploration, parallel approaches, cleanup, fixture creation, documentation, or dashboarding.
  • What should happen near timeout? Write the report first, preserve partial results, list commands, and leave a clean next step.

The important move is that the skill makes these policies explicit before execution. Instead of correcting the agent after it drifts, you front-load the taste and evidence standard into the goal itself.

The eval

Once goal writing becomes a skill, it can be evaluated like software. The core eval is simple: take a past non-trivial engineering session, hide the successful trajectory, expose only the original task prompt, ask different goal-writing skills to write goals, let executor agents run from those goals under the same budget, and judge the resulting artifacts.

The past session is not copied as a script. It is used as a reference for what a good result looked like: which risks mattered, which evidence was valuable, where the agent needed autonomy, and what mistakes wasted time. The candidate goal writer does not get those answers. The judge can use them later to ask whether the new run reached comparable value.

evals/ scanner-optimization.eval.json fixtures/ scanner-optimization/ TASK.md # public prompt visible to candidates REFERENCE_TRAJECTORY.md # hidden lessons from prior good session skills/ baseline-goal-writer/SKILL.md autonomy-goal-writer/SKILL.md systems/ pi-coding-agent-base-instructions.md runs/ 20260523_baseline_gpt55/ GOAL.md EXECUTOR_REPORT.md judge.result.json 20260523_autonomy_gpt55/ GOAL.md EXECUTOR_REPORT.md judge.result.json

This is deliberately practical. The output is not a leaderboard in the abstract. It is a pile of concrete artifacts: the generated goals, the code or reports produced under those goals, the judge scores, the diffs between skills, and the suggested skill change that would improve the next run most.

The harness matrix

The system under test should be explicit. Otherwise you do not know what won. Was it the model? The harness? The base instructions? The goal-writing skill? The time budget? The judge prompt?

A useful manifest records all of it:

{ "coding_harness": "pi-coding-agent", "base_instructions_file": "systems/pi-coding-agent-base-instructions.md", "goal_writing_skill": "skills/autonomy-goal-writer/SKILL.md", "goal_model": "gpt-5.5", "exec_model": "gpt-5.3-codex", "judge_model": "gpt-5.5", "reasoning": "high", "time_limit_seconds": 2700, "token_budget": "unlimited", "goal_phase_tools": "disabled", "execution_phase_tools": "enabled", "judge_phase_tools": "disabled" }

This lets you compare meaningful combinations: the same goal-writing skill across two models, two skills on one model, two harnesses with the same base instructions, or a new base instruction set against the default. The goal is not to declare one universal winner. The goal is to learn which operating contract produces the best work for a class of task.

Judging artifacts

The judge should evaluate artifacts, not vibes. For a performance task, that means correctness preservation, benchmark methodology, before/after speed, and whether the result can be reproduced. For a UI task, that means working interactions, source-backed behavior, visual fidelity, and no screenshot facsimiles pretending to be implementation. For research, it means citations, uncertainty tracking, and decision quality.

Judge dimensions
DimensionQuestionFailure mode
Correctnessdid it preserve required behavior?fast but wrong
Evidencecan a reviewer reproduce the claim?unsupported summary
Autonomydid it make good local decisions?asked too early or thrashed
Tastedid it match project/user standards?technically ok, wrong shape
Reportingis the review surface clear?work exists but is hard to inspect
Leveragewhat skill change improves next run?score without learning

The last row is the most important. A judge result should not only say “88/100.” It should say which prompt or skill change would have improved the score most. That turns every eval run into training data for the operating system around the agent.

Review becomes data

In normal agent work, human review disappears into chat. The human says “this is fake,” “you should have benchmarked,” “open the actual app,” “do not stop at a plan,” or “compare against the original behavior.” The agent fixes the local issue, but the correction is rarely promoted into a reusable rule.

Goal-writing evals create a path for promotion. A review comment can become:

  • a judge criterion,
  • a fixture based on a past failure,
  • a line in the goal-writing skill,
  • a base instruction for the harness,
  • a dashboard field reviewers inspect every run, or
  • a stop rule that preserves evidence before timeout.

This is the compounding loop. The user spends less time steering individual runs and more time reviewing examples. The system converts those reviews into better goals. Better goals produce better autonomous work. Better work creates cleaner artifacts to review.

Why it matters

The ceiling on agentic coding is not only model intelligence. It is task framing, evidence discipline, and feedback capture. A very strong model can still waste 45 minutes if the goal does not say what counts as success. A weaker model can sometimes perform surprisingly well if the goal gives it the right rails.

Goal-writing evals matter because they make this measurable. They answer questions that otherwise become anecdotes:

  • Does a more autonomous goal actually reduce intervention?
  • Does it improve final artifacts or merely produce longer plans?
  • Which model benefits most from stronger goals?
  • Which tasks need stricter evidence instructions?
  • Which human review comments repeat often enough to become policy?

This is software engineering applied to software engineering. You build an eval harness around your own work loop. You version the instructions. You compare outputs. You keep the artifacts. You improve the skill based on failures. The automation target is not a shell command. It is the whole path from intent to trusted result.

How to start

Start small. Pick one past session where the work was non-trivial and the final result was genuinely useful. Extract a public task prompt and a hidden reference note. Write two goal-writing skills: one baseline and one opinionated. Run both under the same harness, model, reasoning level, and time budget. Judge the artifacts. Then change exactly one thing.

# one fixture, two skills, one harness run_eval --eval evals/scanner-optimization.eval.json --skill skills/baseline-goal-writer/SKILL.md --harness pi-coding-agent --model gpt-5.5 --reasoning high --time-limit 45m run_eval --eval evals/scanner-optimization.eval.json --skill skills/autonomy-goal-writer/SKILL.md --harness pi-coding-agent --model gpt-5.5 --reasoning high --time-limit 45m

The first version does not need to be perfect. It needs to be real. Real generated goals. Real executor artifacts. Real judge output. Real failures. Once those exist, improving the system becomes normal engineering: add fixtures, tighten schemas, calibrate judges, build a dashboard, track token usage, and promote repeated review comments into the skill.

That is the broader idea: do not merely use agents to automate code edits. Automate the conditions that make autonomous code work trustworthy. The goal is the interface. The eval is the test. The review is the training signal. The software engineer's job is to connect them into a loop that gets better.