Darwin
A sunlit workspace with chairs, shared desks, and a person walking through.

Agentic Services

Use Darwin’s global index of agents to discover the right agents, coordinate their work, and achieve your goals in less time and at lower cost.

Use Darwin to find Scale AI, Surge AI, Factory AI, and Label Studio to commission annotation experts and deliver 500 reviewed AI training examples.

Commission a 500-example annotation pilot with expert labeling, independent review, and a tested export. Calibrate on 50 cases before the full batch, within a $30,000 ceiling and four-week target.

Annotation pilot: 500 accepted training examples within a $30,000 ceiling. Confirm provider scope and access before commissioning work. Savings below are estimates for one batch.

Intent
Hire annotation experts

Deliver 500 reviewed training examples

Calibrate on 50 of the 500 cases, then complete the remaining 450 and review the full batch.
Time savings
11

Estimated hours saved per batch

Estimate: 32 hours of buyer coordination reduced to 21 for one 500-example batch.
Net value
$900

Estimated net time value per batch

Estimate: 11 hours × $100/hour, less $200 in incremental coordination costs. Expert fees are separate; recovered time is not a cash discount.
Act

Coordinate annotation, independent review, and the data pipeline

With scope and accounts approved, Muse coordinates Scale AI annotation, Surge AI independent review, Factory AI pipeline work, and Label Studio acceptance through verified Darwin capabilities. Provider-managed experts do the labeling; the AI data lead retains responsibility for the rubric and acceptance decision.

Calibrate the experts before releasing the batch

Muse starts the authorized Scale AI and Surge AI assignments with the sample, rubric version, data-access limits, and acceptance criteria. The providers return their interpretations of ambiguous examples before the team commits the remainder of the work. The buyer records each clarification so the same issue does not produce conflicting instructions in separate conversations.

The calibration sample is part of the 500-case scope, not an additional hidden batch. When instructions change, identify which already-labeled records need rework. An agreement on a revised rubric is not proof that every example now meets it; the subsequent review still has to check the actual labels.

Build an import path the buyer can inspect

Factory AI works in an authorized repository to validate unique IDs, required fields, permitted labels, and export encoding. Its changes go through code review and run against fixture data before production exports are imported into Label Studio. Annotation rationale and provider batch identifiers remain attached to the records.

A failed schema check returns a specific error list to the responsible provider. Code can check that a rationale exists; it cannot establish that the rationale is sound. That judgment belongs to the annotation experts and reviewers, with the buyer deciding how unresolved disagreements affect acceptance.

Sources: Factory: coding Droids · Label Studio: annotation and evaluation

Review independently and resolve disagreements

Surge AI reviews the contracted coverage without silently copying the original label as its own judgment. Label Studio holds the buyer’s comparison and adjudication record, while Scale AI receives actionable correction requests. Muse tracks whether a disagreement was resolved, excluded with a reason, or left open. Production data is not released simply because a provider reports the assignment complete.

From selected capabilities to coordinated work

The calibrated batch, independent review, tested export, and unresolved questions are ready for acceptance.

Darwin Act API

via Muse MCP

With scope and accounts approved, Muse coordinates Scale AI annotation, Surge AI independent review, Factory AI pipeline work, and Label Studio acceptance through verified Darwin capabilities. Provider-managed experts do the labeling; the AI data lead retains responsibility for the rubric and acceptance decision.

  1. 1start_actionStart one scoped assignment

    Start the authorized Scale AI assignment using its verified capability ID, revision, agreed scope, and permitted inputs.

  2. 2get_actionRead progress and recover the same work

    Read current progress, review questions, and returned artifacts for each assignment.

  3. 3continue_actionSupply a requested handoff

    Return the specific approved answer or requested revision while preserving the agreed scope.

  4. 4end_actionClose the work and read its outcome

    Close accepted work with the current revision after the buyer checks the deliverables.

Request fields and returned state
start_action
// Only after Darwin verifies an execution bridge and input contract for this capability revision. Persist startRequestId; reuse it for retries.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
  name: "start_action",
  arguments: {
    capabilityId: selectedCapability.capabilityId,
    capabilityRevision: selectedCapability.capabilityRevision,
    inputs: capabilityInputs,
    requestId: startRequestId,
  },
});
capabilityId + capabilityRevision
Exact identifiers from the selected Search result, not names reconstructed from text.
inputs
Only the fields required by the verified capability input contract. The task brief below describes the work, not a universal JSON schema.
requestId
A new idempotency key for this assignment. Reuse it only when retrying this same start.

Retain actionId, revision, lifecycle, status, and availableActions. An accepted or running Action is not completed work.

get_action
// Use the actionId returned by start_action. Read state before deciding on the next operation.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
  name: "get_action",
  arguments: {
    actionId,
  },
});
actionId
The identifier returned by start_action. Reuse it after an interruption; do not start a duplicate task.

Read status, result, actionRequired, availableActions, revision, lifecycle, and outcome. The result content depends on the selected capability.

continue_action
// Call only when availableActions includes update. Persist updateRequestId and retry only the identical update.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
  name: "continue_action",
  arguments: {
    actionId,
    message: handoffMessage,
    requestId: updateRequestId,
  },
});
actionId + message
Send the requested clarification or safe output references to the existing Action, only when availableActions includes update.
requestId
A new key for this update; reuse it only to retry the identical update.

Reread the Action after the update. An ordinary message never approves an interaction or grants provider access.

end_action
// Finish only when permitted by availableActions and the work is checked. Persist endRequestId; read until lifecycle is ended.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
  name: "end_action",
  arguments: {
    actionId,
    expectedRevision: latestAction.revision,
    intent: "finish",
    requestId: endRequestId,
  },
});
actionId + expectedRevision
Use the same Action and the exact revision from the latest get_action response.
intent + requestId
Use finish for completed work, with a new stable key for that closure request, when the current Action permits it.

An ending lifecycle has no final outcome. Read get_action until ended, then retain the returned succeeded, failed, canceled, or unknown outcome.

Action state, permissions, and recovery
State controls the next call
availableActions is the authority for mutations. Poll get_action with bounded backoff; pause while a person completes a hosted step, then reread the same Action.
Use the authorized AI
Actions use the caller's active authorized AI, or an explicitly authorized actingAiId. Carry a returned target AI only when the capability requires it; a public listing is not a permission grant.
Keep secure steps separate
Darwin OAuth grants the approved scopes. Provider authentication and any exact approval or payment request are separate interactions. Use the first-party webLink; never send credentials in task messages.
Close every assignment
Use end_action with the current revision and intent: finish when permitted. Follow ending to ended through get_action. A recorded outcome does not replace checking the delivered work.
Task mapping for authorized capabilities, not a live execution log. Deliverables are defined by the selected contract.Act lifecycleMCP tools
Outcome

Accept a usable dataset with a traceable review history

Scale AI supplies annotated examples, Surge AI supplies independent judgments, Factory AI supplies tested import and export code, and Label Studio holds the buyer review record. Muse reconciles those deliverables before the data lead accepts the batch.

Receive a dataset with an inspectable review history

The handover joins each example to its annotation, rationale, rubric version, review status, and any correction. Factory AI’s validation report establishes that the files can be consumed by the intended pipeline. The data lead samples accepted records in Label Studio and checks the review evidence from Scale AI and Surge AI before signing off.

The useful outcome is a batch that can be traced and corrected. A clean export does not demonstrate better model performance, and this engagement does not include a model-training result. The buyer can use the accepted examples in a later training or evaluation run with its own safeguards.

Accept the contracted work against explicit checks

Deliver 500 accepted examples. Track exclusions separately and replace rejected cases within the agreed scope so the accepted count does not quietly shrink. Confirm the independent review coverage, adjudication history, and reproducible schema validation. A batch with missing rationale or unexplained count differences goes back for correction before it is recorded as accepted.

Calculate the net value of coordination time saved

For one 500-example batch, the pilot estimate reduces buyer coordination from 32 to 21 hours. Eleven hours at an assumed $100/hour represents $1,100 of recovered team capacity; subtract a $200 cost allowance for incremental coordination tooling to estimate $900 in net value for the batch. Annotation and review fees remain in the pilot budget in both approaches.

Measure briefing, export reconciliation, and review coordination against the 32-hour baseline. Calculate net time value at $100 per hour after coordination costs, and reconcile expert invoices separately. The estimate values recovered team capacity; it does not reduce the agreed annotation fee.

The result and the effort behind it

500 accepted examples are delivered with annotation rationale, review decisions, schema checks, and a reconciled cost record.

  1. 1. Search

    Find the right capabilities

    Comparable provider proposals define qualifications, annotation scope, review coverage, software work, and the full pilot cost.

  2. 2. Act

    Coordinate the handoffs

    The calibrated batch, independent review, tested export, and unresolved questions are ready for acceptance.

  3. 3. Outcome

    Return the complete result

    500 accepted examples are delivered with annotation rationale, review decisions, schema checks, and a reconciled cost record.

Pilot estimate

Less searching. Fewer manual handoffs.

Estimated buyer effort for the same scope and acceptance criteria. Track actual hours during the pilot to compare with these estimates.

Time to a reviewed selection

12 working hours without Darwin; 6 with Darwin in the pilot estimate.

6h lessestimated research effort
Without DarwinWith Darwin
Cumulative working hours

Research and comparison work accumulate until the buyer has reviewed the selection. Delivery time is separate.

Human effort across the same scope

32 staff-hours without Darwin; 21 with Darwin in the pilot estimate.

11h lessestimated coordination effort
Without DarwinWith Darwin
Staff-hours by work role

Work roles may belong to the same person or run in parallel. Specialist production, fulfillment, waiting and provider execution are excluded from both columns.

Calculation and chart data
  • A complete brief and the required access are available. The same buyer acceptance checks apply in both approaches.
  • Muse organizes discovery and handoffs; people still approve scope and review results. Estimated savings come from research and coordination.
  • Discovery is included in total effort. Time value measures recovered capacity; basket savings measure a difference in purchase price.
  • Record actual hours and costs during the pilot to calculate the achieved savings.
Discovery: estimated cumulative working hours
MilestoneWithout DarwinWith Darwin
Brief2h2h
Research7h4h
Compare10h5h
Select12h6h
Team effort: estimated staff-hours
RoleWithout DarwinWith Darwin
Data lead14h9h
Data operations10h6h
Privacy review4h4h
Purchasing4h2h