September 23 / Part 1

From Intent to Pull Request

Building a Remote Engineering Harness

Who's talking

Bharadwaj
Pendyala

Lead Member of Technical Staff at Salesforce. I work on enterprise workflow systems and the tooling around agents.

I use coding agents in my own development workflow.

Part 1 / Before and after

What coding agents changed

Before coding agents

Two to three weeks per cycle

Requirements, design, build, test, deploy. One sequence, and the idea only met reality at the end of it.

After coding agents

Zero to one got cheap

Validate the idea, make a mock, hand it to an agent, put it in front of beta users. Then run the loop again.

Not just a faster lifecycle. The decision-making loop got shorter, and the work moved out of isolated development environments into shared, governed workspaces.

Part 1 / Zero to one

Zero to one in hours

The zero to one loop A six stage ring running clockwise: align with the stakeholder, turn the meeting transcript into a specification, refine it with constraints and impact analysis, let the agents execute, run an adversarial review and a human test, then ship to a beta group whose feedback starts the loop again. 0101 · ALIGNSTAKEHOLDERTalk to the stakeholderWhat breaks today, why it matters now,and what changes once it works. 0202 · CAPTUREAGENTTranscript becomes a specThe meeting recording goes to the agent,which drafts the specification. 0303 · CONSTRAINENGINEERRefine and check impactAdd the constraints. Name every otherfeature this change touches. 0404 · EXECUTEAGENTSAgents run the workDependent steps in order. Independentsteps at the same time. 0505 · REVIEWAGENT, HUMANAdversarial reviewA second agent argues against the diff.Then a human tests what matters. 0606 · RELEASEBETA USERSShip to a beta groupA small group, a small blast radius.Their reaction is the next input. One loop Hours not weeks

The stakeholder does no new work. It is the conversation they already have, captured once.

Part 1 / What the loop needs

Autonomous from idea to draft PR

Where the run is autonomousThree bands. A human band on the left, where the stakeholder answers the questions and the engineer sets the constraints. A wide autonomous band in the middle, where the agent drafts the spec, implements, checks and records the behavior, reviews the diff against itself, and opens the draft pull request. A human band on the right, where the stakeholder accepts the behavior and the engineer reviews and merges.HumanAutonomousHumanThe agreed requestStakeholder answers the questions.Engineer sets the constraints.No person in the middleApproved specfrom the conversationImplementand repair in limitsCheck and recordthe working behaviorAdversarial reviewthen the draft PROne approved spec. A failure stops or uses a bounded repair.The judgmentStakeholder accepts the behavior.Engineer reviews and merges.

Autonomy ends at the draft pull request. An engineer still owns the merge.

Part 1 / Where the loop stalls

Every change queues behind an engineer

Where the work queues today, and where it queues with the harnessTwo lanes compared column by column. Today: anyone with a small ask, the request joins the engineer's queue behind work that needs their judgment, an engineer builds it eventually, and a change arrives days later. With the harness: anyone starts from where they already work, the harness starts the run in a remote environment, agents implement and verify inside the boundary, and a draft pull request or a direct answer comes back. Complex features still start with an engineer.TodayWHO STARTSAnyone with a small askA filter that does not exist yet, afield that should survive a reload.WHAT HAPPENSIt joins the engineer's queueBehind the work that actuallyneeds their judgment.WHO BUILDSAn engineer, eventuallyContext switch, branch, localrun, then a pull request.WHAT COMES BACKA change, days laterIf it was still worth doing bythe time it reached the top.With the harnessWHO STARTSAnyone, where they workFrom Slack, a meeting, or aticket. No repository, no terminal.WHAT HAPPENSThe harness starts the runA remote environment and theconstraints the engineer set.WHO BUILDSAgents, inside the boundaryImplement, verify, record, andargue against the diff.WHAT COMES BACKA draft PR with the evidenceA short video to accept and adiff to review, on one commit.Complex features still start with an engineer. This path is for small fixes, small features, and questions about the code. The demo shows one small feature.

The bottleneck is the handoff, not the building. So remove the handoff.

Part 1 / The idea

A remote harness

Take an agreed request, run it on a machine nobody had to configure, inside constraints an engineer set once, and hand back work that is ready to review.

Part 1 / 08

What the harness has to provide

Proposed first-version requirements

  • A place to make the request where the conversation already happens.
  • A remote environment, so nobody configures anything locally.
  • Constraints and permissions an engineer sets once, for the repository.
  • Evidence for two readers: a recording to accept, a diff to review.
  • A stopping point for when the answer is not in the repository.
Part 1 / 09

Start with a question, end with a reviewable change

  • Ask what the application does today.
  • Discuss a change and clarify its behavior in the same conversation.
  • Review a spec revision and explicitly approve the build.
  • Actions implements, checks, records, and opens a draft PR.
  • The stakeholder accepts the behavior. An engineer reviews and merges.
Part 1 / 10

“Is an Important label implemented?”

  • Start with the existing to-do app.
  • The agent reads the pinned source and answers with file references.
  • Then ask: “Let’s add a star marker that survives reload.”
Part 1 / 11

Questions the stakeholder can answer

  • An on/off star, or several priority levels?
  • A marker only, or also a filter or sorting?
  • Should the marker survive a reload?

Our example chooses a persistent star marker. No filter or sorting.

Part 1 / 12

The conversation stays open

Stakeholder: “Let’s add Important.”

Harness: “A star marker only, or should it filter or reorder tasks?”

Stakeholder: “A star only. Keep the order.”

Harness: “Should it survive reload?”

Stakeholder: “Yes. Keep completion working as it does.”

Part 1 / 13

One requirement becomes an observable check

Illustrative Playwright check; selectors follow the generated UI

important-marker.spec.tsworked example
test('Important survives reload', async ({ page }) => {
  const task = page.getByRole('listitem').filter({ hasText: 'Task A' });
  await task.getByRole('button', { name: 'Mark important', exact: true }).click();
  await page.reload();
  await expect(task.getByRole('button', { name: 'Remove important', exact: true }))
    .toHaveAttribute('aria-pressed', 'true');
});

Mark the task before reload. Assert the stored marker afterward.

Part 1 / 14

Approve this exact specification

  • Toggle an Important star on an individual task.
  • Persist it across reload; existing tasks start unmarked.
  • Keep completion and list order unchanged.
  • Do not add a filter, sorting, or priority levels.

The approval button names a spec revision. Editing the requirements invalidates the earlier draft.

Part 1 / 15

Chat stays warm, builds run on Actions

Conversation, approval, execution and reviewStreamlit exchanges messages with a persistent Node service running Claude Agent SDK. It reads a pinned source commit and saves conversation history and spec revisions. Explicit human approval writes a run record and dispatches GitHub Actions. The worker checks and records a candidate; a separate publisher opens the draft pull request. Status and the PR link return to chat.Streamlit chatAsk and clarify.Review spec revision.Approve the build.Receive PR + video.Persistent conversation serviceNode + Claude Agent SDKRead-only tools over a pinned Git commitTranscript + session ID + draft revisionsOnly approval dispatches execution.Durable storage survives reconnects.GitHub ActionsGate → implement → check → recordSeparate review and bounded repairPublisher opens the draft PR.harness-state stores the run.approved specstatus + PR

Open the architecture reference map

DEMO

Part 1 / 17

What changed, and what stayed

Conversation

  • Streamlit sends messages to the service.
  • Claude Agent SDK keeps the conversation alive.
  • Repository tools read a pinned commit.
  • A draft becomes executable only after human approval.

Execution

  • Run JSON lives on harness-state.
  • Actions runs implementation and checks.
  • Playwright records the candidate.
  • A separate publisher holds the app write token.
Part 1 / Observed

The harness cannot live in the repository it changes

Observed while building this, not a worked example

  • The first version kept the coordinator inside the application repository.
  • Every run checks out the base commit before the agent starts.
  • That checkout reverted the coordinator to a version without the command that was running.
  • The run record was in the same repository, so it went too.

Durable and ephemeral is not a drawing convention. Ignore it and a run deletes the thing running it. The harness is now a separate repository, and the record is a branch it never checks out.

Part 1 / 19

What starts an Actions runner

OperationWhere it runsResult
Question or follow-upPersistent SDK serviceStreamed answer; no workflow
Clarify and revise a specSame SDK conversationNew draft revision
Approve and buildexecute.yml, mode executeCandidate, checks, video, review
Repair a failed candidateexecute.yml, mode repairAnother check, or a retained failure

The gate checks the current record. The publisher runs after review. The router records the outcome.

Part 1 / 20

Preparing a repository for the harness

The engineer supplies:

  • Four scripts: setup, start, check, record.
  • One adapter file naming them, harness.yml.
  • Seed data, and Playwright as the check runner.
  • A token scoped to one repository, held by one job.

The stakeholder supplies the desired behavior.

The repository adapter makes setup explicit

todo-app/harness.yml documents the app contract

yaml
setup: ./scripts/setup-app
start: ./scripts/start-app
ready: http://127.0.0.1:3000/health
checks:
  - ./scripts/check-app
record: ./scripts/record-journey
artifacts: /run-output

The journey name comes from the spec. The current worker calls these script paths directly; it does not yet load arbitrary adapters.

The same adapter, a real repository

Current app contract / illustrative service adaptation

to-do appyaml
setup: ./scripts/setup-app
start: ./scripts/start-app
ready: http://127.0.0.1:3000/health
checks:
  - ./scripts/check-app
record: ./scripts/record-journey
artifacts: /run-output
a service with a databaseyaml
setup: npm ci && docker compose up -d db && npm run migrate
start: docker compose up api web
ready: http://127.0.0.1:8080/healthz
checks:
  - npm run test:integration
record: ./scripts/record-journey checkout
artifacts: /run-output

The contract can stay small. A generic adapter loader is still future work.

Part 1 / 23

Give each part only the tools it needs

  • The chat agent gets repository reads and a draft-spec tool. No shell, edits, or GitHub publishing credentials.
  • The service writes harness run records and dispatches approved builds.
  • The worker edits and checks its disposable checkout.
  • The separate publisher pushes the candidate and opens a draft PR.

An engineer owns merge. Token scope and branch rules must enforce the intended repository policy.

Part 1 / 24

Implementation, checks, and the recording

  1. Implement the agreed behaviour on a branch from the base commit.
  2. Run the checks. A failure costs a repair from the budget.
  3. Drive the feature in a real browser and record it.
  4. A second agent argues against the diff.
  5. Only then does anything reach a pull request.

The video, the checks, and the reviewer all refer to the same candidate commit.

Part 1 / 25

Conversation state and execution evidence

WhatWhereIdentity
Messages, SDK transcript, draft revisionsService durable volumeConversation ID + source commit
Approved spec and build stateharness-state/runs/<id>.jsonRun ID + approved spec hash
Candidate and evidenceApp branch, Actions artifacts, draft PRCandidate commit + run ID

Prepared example: todo-app PR #4, created by the earlier intake flow.

Part 1 / 26

A failure decides what happens next

Observed / the checks passed, the recording did not

  • The checks passed. The implementation was fine.
  • The recording failed: one locator matched six elements.
  • The coordinator read the state and dispatched a repair.
  • The repair was told to fix the journey only, never the implementation.
  • Budget spent and the run stops, with the log kept and no pull request opened.
Part 1 / 27

The recording crosses the persistence boundary

  • Show the baseline task list.
  • Mark Task A Important, then reload.
  • Show that the star is still set.
  • Complete Task A; its star remains set.
  • Remove the star and show that list order did not change.

The generated recording is attached to the draft PR.

Part 1 / Review handoff

A review package for two audiences

Above the line: the request, what it does, the video, the agreed behaviour.

Below the line: the review focus, the reviewer’s argument, the checks, the diff.

One pull request, one candidate commit, two readers who never have to read each other’s half.

Part 1 / 29

The body of the draft pull request

draft PR bodystructure
### What it does
Add a persistent Important star without changing list order.

[Generated Playwright video]

### The behaviour that was agreed
- The star survives reload. Completion still works.

### Review focus
Existing tasks migrate without losing their data.

### Checks
Command output and candidate commit.
Part 1 / 30

Carry source findings into the build

  • The baseline task table contains id, title, and done.
  • Adding a persistent marker needs a schema change that handles existing rows.
  • The approved run carries these findings and the pinned source commit into the worker.
Part 1 / Review handoff

Acceptance and engineering review

The stakeholder confirms: “This is the behavior I wanted.”

The engineer reviews the implementation, checks, and architecture fit.

The engineer owns merge through the repository's normal process.

Part 1 / The next loop

The loop closes with the next request

The harness stops at the draft pull request. Release stays on the path your team already uses.

  • A merged change goes to a small group before it goes to everyone.
  • What that group does with it is the evidence.
  • A confusing result becomes the next request, not a bug report.
  • That request enters at the top of the loop, and the run happens again.
Part 1 / 33

When a run stops

  • An unresolved product question returns to the stakeholder.
  • A new architecture decision goes to an engineer.
  • A check, a recording, or a reviewer can send a run back for one repair.
  • When the budget is spent the run stops, keeps its reason and its artifacts, and opens nothing.
Part 1 / 34

The starter people can run

  • remote-harness: persistent SDK chat, Streamlit, approval and Actions.
  • todo-app: baseline app, instructions, checks and recording scripts.
  • Local and container setup with persistent chat storage.
  • A prepared draft PR and an architecture reference map.

Adapting the starter to your application

Proposed starter adaptation

PieceTo-do exampleYour repository
SetupTask data and app runtimeDependencies and services
InstructionsTask-list conventionsArchitecture and contribution rules
ChecksImportant flag and reloadYour acceptance behavior
RecordingMark, reload, completeA short user journey
PublicationVideo and draft PRRepository and reviewer access

The first milestone is one complete run in a fresh environment.

Building the contribution workflow

Engineers can build the systems that let other roles contribute.

  • Today: one remote request-to-PR workflow.
  • Part 2: making the harness more effective and reliable.
  • Part 3: software factories across workflows and teams.

Questions

Questions or ideas?

Thank you.

Arrow keys to navigate / N notes / F fullscreen