From Intent to Pull Request
Building a Remote Engineering Harness
Bharadwaj Pendyala
Bharadwaj
Pendyala
Lead Member of Technical Staff at Salesforce. I work on enterprise workflow systems and the tooling around agents.
I use coding agents in my own development workflow.
- Slidesbharad.dev/talks
- LinkedInin/bharadwajpendyala
- GitHubbharadwaj-pendyala
- X@bharadwaj_py
What coding agents changed
Two to three weeks per cycle
Requirements, design, build, test, deploy. One sequence, and the idea only met reality at the end of it.
Zero to one got cheap
Validate the idea, make a mock, hand it to an agent, put it in front of beta users. Then run the loop again.
Not just a faster lifecycle. The decision-making loop got shorter, and the work moved out of isolated development environments into shared, governed workspaces.
Zero to one in hours
The stakeholder does no new work. It is the conversation they already have, captured once.
Autonomous from idea to draft PR
Autonomy ends at the draft pull request. An engineer still owns the merge.
Every change queues behind an engineer
The bottleneck is the handoff, not the building. So remove the handoff.
A remote harness
Take an agreed request, run it on a machine nobody had to configure, inside constraints an engineer set once, and hand back work that is ready to review.
What the harness has to provide
Proposed first-version requirements
- A place to make the request where the conversation already happens.
- A remote environment, so nobody configures anything locally.
- Constraints and permissions an engineer sets once, for the repository.
- Evidence for two readers: a recording to accept, a diff to review.
- A stopping point for when the answer is not in the repository.
Start with a question, end with a reviewable change
- Ask what the application does today.
- Discuss a change and clarify its behavior in the same conversation.
- Review a spec revision and explicitly approve the build.
- Actions implements, checks, records, and opens a draft PR.
- The stakeholder accepts the behavior. An engineer reviews and merges.
“Is an Important label implemented?”
- Start with the existing to-do app.
- The agent reads the pinned source and answers with file references.
- Then ask: “Let’s add a star marker that survives reload.”
Questions the stakeholder can answer
- An on/off star, or several priority levels?
- A marker only, or also a filter or sorting?
- Should the marker survive a reload?
Our example chooses a persistent star marker. No filter or sorting.
The conversation stays open
Stakeholder: “Let’s add Important.”
Harness: “A star marker only, or should it filter or reorder tasks?”
Stakeholder: “A star only. Keep the order.”
Harness: “Should it survive reload?”
Stakeholder: “Yes. Keep completion working as it does.”
One requirement becomes an observable check
Illustrative Playwright check; selectors follow the generated UI
test('Important survives reload', async ({ page }) => {
const task = page.getByRole('listitem').filter({ hasText: 'Task A' });
await task.getByRole('button', { name: 'Mark important', exact: true }).click();
await page.reload();
await expect(task.getByRole('button', { name: 'Remove important', exact: true }))
.toHaveAttribute('aria-pressed', 'true');
});Mark the task before reload. Assert the stored marker afterward.
Approve this exact specification
- Toggle an Important star on an individual task.
- Persist it across reload; existing tasks start unmarked.
- Keep completion and list order unchanged.
- Do not add a filter, sorting, or priority levels.
The approval button names a spec revision. Editing the requirements invalidates the earlier draft.
Chat stays warm, builds run on Actions
DEMO
What changed, and what stayed
Conversation
- Streamlit sends messages to the service.
- Claude Agent SDK keeps the conversation alive.
- Repository tools read a pinned commit.
- A draft becomes executable only after human approval.
Execution
- Run JSON lives on
harness-state. - Actions runs implementation and checks.
- Playwright records the candidate.
- A separate publisher holds the app write token.
The harness cannot live in the repository it changes
Observed while building this, not a worked example
- The first version kept the coordinator inside the application repository.
- Every run checks out the base commit before the agent starts.
- That checkout reverted the coordinator to a version without the command that was running.
- The run record was in the same repository, so it went too.
Durable and ephemeral is not a drawing convention. Ignore it and a run deletes the thing running it. The harness is now a separate repository, and the record is a branch it never checks out.
What starts an Actions runner
| Operation | Where it runs | Result |
|---|---|---|
| Question or follow-up | Persistent SDK service | Streamed answer; no workflow |
| Clarify and revise a spec | Same SDK conversation | New draft revision |
| Approve and build | execute.yml, mode execute | Candidate, checks, video, review |
| Repair a failed candidate | execute.yml, mode repair | Another check, or a retained failure |
The gate checks the current record. The publisher runs after review. The router records the outcome.
Preparing a repository for the harness
The engineer supplies:
- Four scripts: setup, start, check, record.
- One adapter file naming them,
harness.yml. - Seed data, and Playwright as the check runner.
- A token scoped to one repository, held by one job.
The stakeholder supplies the desired behavior.
The repository adapter makes setup explicit
todo-app/harness.yml documents the app contract
setup: ./scripts/setup-app start: ./scripts/start-app ready: http://127.0.0.1:3000/health checks: - ./scripts/check-app record: ./scripts/record-journey artifacts: /run-output
The journey name comes from the spec. The current worker calls these script paths directly; it does not yet load arbitrary adapters.
The same adapter, a real repository
Current app contract / illustrative service adaptation
setup: ./scripts/setup-app start: ./scripts/start-app ready: http://127.0.0.1:3000/health checks: - ./scripts/check-app record: ./scripts/record-journey artifacts: /run-output
setup: npm ci && docker compose up -d db && npm run migrate start: docker compose up api web ready: http://127.0.0.1:8080/healthz checks: - npm run test:integration record: ./scripts/record-journey checkout artifacts: /run-output
The contract can stay small. A generic adapter loader is still future work.
Give each part only the tools it needs
- The chat agent gets repository reads and a draft-spec tool. No shell, edits, or GitHub publishing credentials.
- The service writes harness run records and dispatches approved builds.
- The worker edits and checks its disposable checkout.
- The separate publisher pushes the candidate and opens a draft PR.
An engineer owns merge. Token scope and branch rules must enforce the intended repository policy.
Implementation, checks, and the recording
- Implement the agreed behaviour on a branch from the base commit.
- Run the checks. A failure costs a repair from the budget.
- Drive the feature in a real browser and record it.
- A second agent argues against the diff.
- Only then does anything reach a pull request.
The video, the checks, and the reviewer all refer to the same candidate commit.
Conversation state and execution evidence
| What | Where | Identity |
|---|---|---|
| Messages, SDK transcript, draft revisions | Service durable volume | Conversation ID + source commit |
| Approved spec and build state | harness-state/runs/<id>.json | Run ID + approved spec hash |
| Candidate and evidence | App branch, Actions artifacts, draft PR | Candidate commit + run ID |
Prepared example: todo-app PR #4, created by the earlier intake flow.
A failure decides what happens next
Observed / the checks passed, the recording did not
- The checks passed. The implementation was fine.
- The recording failed: one locator matched six elements.
- The coordinator read the state and dispatched a repair.
- The repair was told to fix the journey only, never the implementation.
- Budget spent and the run stops, with the log kept and no pull request opened.
The recording crosses the persistence boundary
- Show the baseline task list.
- Mark Task A Important, then reload.
- Show that the star is still set.
- Complete Task A; its star remains set.
- Remove the star and show that list order did not change.
The generated recording is attached to the draft PR.
A review package for two audiences
Above the line: the request, what it does, the video, the agreed behaviour.
Below the line: the review focus, the reviewer’s argument, the checks, the diff.
One pull request, one candidate commit, two readers who never have to read each other’s half.
The body of the draft pull request
### What it does Add a persistent Important star without changing list order. [Generated Playwright video] ### The behaviour that was agreed - The star survives reload. Completion still works. ### Review focus Existing tasks migrate without losing their data. ### Checks Command output and candidate commit.
Carry source findings into the build
- The baseline task table contains
id,title, anddone. - Adding a persistent marker needs a schema change that handles existing rows.
- The approved run carries these findings and the pinned source commit into the worker.
Acceptance and engineering review
The stakeholder confirms: “This is the behavior I wanted.”
The engineer reviews the implementation, checks, and architecture fit.
The engineer owns merge through the repository's normal process.
The loop closes with the next request
The harness stops at the draft pull request. Release stays on the path your team already uses.
- A merged change goes to a small group before it goes to everyone.
- What that group does with it is the evidence.
- A confusing result becomes the next request, not a bug report.
- That request enters at the top of the loop, and the run happens again.
When a run stops
- An unresolved product question returns to the stakeholder.
- A new architecture decision goes to an engineer.
- A check, a recording, or a reviewer can send a run back for one repair.
- When the budget is spent the run stops, keeps its reason and its artifacts, and opens nothing.
The starter people can run
- remote-harness: persistent SDK chat, Streamlit, approval and Actions.
- todo-app: baseline app, instructions, checks and recording scripts.
- Local and container setup with persistent chat storage.
- A prepared draft PR and an architecture reference map.
Adapting the starter to your application
Proposed starter adaptation
| Piece | To-do example | Your repository |
|---|---|---|
| Setup | Task data and app runtime | Dependencies and services |
| Instructions | Task-list conventions | Architecture and contribution rules |
| Checks | Important flag and reload | Your acceptance behavior |
| Recording | Mark, reload, complete | A short user journey |
| Publication | Video and draft PR | Repository and reviewer access |
The first milestone is one complete run in a fresh environment.
Engineers can build the systems that let other roles contribute.
- Today: one remote request-to-PR workflow.
- Part 2: making the harness more effective and reliable.
- Part 3: software factories across workflows and teams.
💬 Questions