Skip to main content

Private beta docs

Run Flock reviews from developer workflows

Choose a target URL, personas, and a user goal. Review the findings, inspect available evidence, and prepare fixes for your team or coding agent.

Connect through MCP or the GitHub App. The CLI runs from a Flock runtime checkout, available by request during the beta.

GitHub App

Flock GitHub App

Review supported PR previews with a Preview Gate check and summary. Inline comments are added when findings map to changed code.

Review outputs

  • Preview Gate check
  • PR summary
  • Mapped inline findings
View GitHub App docs
MCP

Flock MCP Server

Connect a supported AI client to recommend and estimate reviews, launch approved tests, and retrieve results through tools.

Review outputs

  • 14 tools
  • Evidence images
  • Repair prompt retrieval
View MCP docs
CLI

Flock Synthetic CLI

Run reviews from a Flock runtime checkout, available by request during the beta.

Review outputs

  • output.json
  • report.md
  • engineering_events.json
View CLI docs

Quickstart

Run one review against a preview URL

CLI example

Replace the example URL with your target. This example uses the bundled skeptical-evaluator persona and runs from a Flock runtime checkout.

npm run tester -- \
  --url https://preview.example.com/signup \
  --persona skeptical-evaluator \
  --goal "Create a shared workspace and invite a teammate" \
  --success-signals '[{"type":"text_present","value":"Invitation sent"}]'

Outputs written

  • output.json for structured findings, evidence, and fix guidance.
  • report.md for a human-readable review summary.
  • engineering_events.json for console, network, and runtime signals.
  • timeline.jsonl for step-by-step replay context.

Read the result and its evidence status

Abbreviated excerpts; omitted fields are not shown. The output and batch summary come from an internal replay evaluation of a synthetic checkout page. Its finding is unverified, not a confirmed defect. The GitHub check is from a separate completed review.

output.json excerpt

{
  "schema_version": "2.0.0",
  "persona_id": "novice-user",
  "page_url": "https://example.test/checkout",
  "perception_scope": "screenshot_only",
  "friction_events": [
    {
      "event_id": "event-1",
      "severity": 2,
      "summary": "The Checkout button is very low-contrast, so the main action is hard to read at a glance.",
      "confidence": 0.5,
      "evidence": {
        "screenshot_ref": "artifacts/screenshot.png",
        "page_url": "https://example.test/checkout",
        "timestamp": null,
        "dom_ref": null,
        "har_ref": null,
        "location": {
          "region": "center",
          "bbox": null
        }
      },
      "verification": {
        "status": "unverified",
        "failure_tags": [
          "webmarker_unavailable",
          "dom_snapshot_unavailable"
        ]
      }
    }
  ]
}

summary.md excerpt

- Status: completed
- Outcome: succeeded
- Runs attempted: 1
- Runs completed: 1
- Runs failed: 0
- Runs summarized: 1

GitHub check excerpt

{
  "name": "Flock CI / Preview Gate",
  "status": "completed",
  "conclusion": "success",
  "output": {
    "title": "First PR run captured current UX context",
    "summary": "Verdict: pass\nRun outcome: pr_local_baseline_seeded\nall target baselines resolved\nBlocking findings on first PR run: 0\nWarning findings on first PR run: 0\nNot observed in the latest comparison: 0\nStill present from baseline: 0"
  }
}

Quality bar

Review findings with their evidence

Check the evidence status

Read each result’s verification and availability information. Missing evidence is not a clean result.

Keep observations distinct

Directional observations provide context. They are separate from actionable findings and do not block the GitHub gate.

Review before changing code

Inspect the available evidence and repair guidance, then verify your change with a focused follow-up review.