Use Darwin to find Scale AI, Surge AI, Factory AI, and Label Studio to commission annotation experts and deliver 500 reviewed AI training examples.
Commission a 500-example annotation pilot with expert labeling, independent review, and a tested export. Calibrate on 50 cases before the full batch, within a $30,000 ceiling and four-week target.
Annotation pilot: 500 accepted training examples within a $30,000 ceiling. Confirm provider scope and access before commissioning work. Savings below are estimates for one batch.
Deliver 500 reviewed training examples
Estimated hours saved per batch
Estimated net time value per batch
Find the experts and tools for 500 reviewed training examples
Muse uses Darwin Search to compare Scale AI and Surge AI expert services, Factory AI implementation support, and Label Studio review tooling against one annotation brief. The provider proposals must specify who labels, who reviews, and how accepted examples become a usable export.
Finding annotators is only the start. The team needs relevant expertise, clear labeling instructions, independent review, and an export that its training pipeline can actually use.
Find annotation expertise that matches the task
The project begins with 500 de-identified customer-support examples that need intent labels, answer-quality judgments, and short rationales. The data lead specifies the domain knowledge required and prepares a 50-case calibration sample containing both routine and ambiguous cases. Scale AI and Surge AI are assessed on their proposed expertise and review process, not simply on the number of workers they can supply.
The buyer commissions a managed service rather than assuming access to named individual freelancers. Agree reviewer qualifications, turnaround, disagreement handling, and what counts as an accepted example. Calibration is part of the paid scope; production begins only once the instructions are understandable and the parties agree how unclear cases will be escalated.
Sources: Scale AI: expert data · Surge AI: expert workforce
Give people and software distinct responsibilities
Scale AI is the proposed primary annotation service and Surge AI the independent review service. This is a division to negotiate, not a claim that either company automatically integrates with the other. Factory AI writes the manifest checks, conversion scripts, and export tests in the buyer’s repository. Label Studio gives the buyer a place to inspect returned labels and resolve review findings.
Providers may require their own production tools. In that case, the contract specifies portable exports that the buyer can import into Label Studio; the scenario does not require external annotators to work inside the buyer’s instance. Stable record IDs and a versioned schema connect the outputs without inventing a shared provider workflow.
Expert judgment, pipeline code, and buyer acceptance have separate owners.
Scale AI: Source annotation experts. Return labels and rationales against the approved rubric, with stable record IDs.
Surge AI: Provide independent expert review. Return independent judgments and actionable corrections for the contracted review coverage.
Factory AI: Build the annotation data pipeline. Validate schema, identifiers, and export counts in the authorized repository.
Label Studio: Review labels and adjudications. Keep the buyer’s inspection, disagreement resolution, and acceptance record together.
Sources: Factory: coding Droids · Label Studio: annotation and evaluation
Set the budget for a reviewed, usable dataset
The $30,000 ceiling includes calibration, annotation, independent review, pipeline work, and required tooling. Quotes must separate those costs and explain rework charges. The four-week target remains a requested schedule until the providers accept it. Hiring more annotators is useful only when the review process can keep up with their output.
From the search brief to a reviewed selection
Comparable provider proposals define qualifications, annotation scope, review coverage, software work, and the full pilot cost.
Darwin Search API
One brief becomes a capability selection
Muse uses Darwin Search to compare Scale AI and Surge AI expert services, Factory AI implementation support, and Label Studio review tooling against one annotation brief. The provider proposals must specify who labels, who reviews, and how accepted examples become a usable export.
search requestFind expert annotation services through Scale AI and Surge AI, Factory AI coding support, and Label Studio review tooling for 500 de-identified support examples. Calibrate on 50 cases and deliver reviewed labels, rationales, and a validated export within a $30,000 ceiling and four-week target.
query contains this brief; limit: 10 bounds the result set. $30,000 pilot ceiling; actual provider quotes establish cost, scope, and availability.
Request example
// Search is read-only. Preserve the selected capability identifiers and revision from the result.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
name: "search",
arguments: {
query: "Find expert annotation services through Scale AI and Surge AI, Factory AI coding support, and Label Studio review tooling for 500 de-identified support examples. Calibrate on 50 cases and deliver reviewed labels, rationales, and a validated export within a $30,000 ceiling and four-week target.",
limit: 10,
},
});- Keep the exact selection
Preserve
aiId,capabilityId, andcapabilityRevisionfor each chosen owner. Keep a returnedsearchAttributionIdso the assignment stays linked to its discovery. - Review the provider scopes before commissioning work
Annotation pilot: 500 accepted training examples within a $30,000 ceiling. Confirm provider scope and access before commissioning work. Savings below are estimates for one batch.
- Carry the selection into Act
Darwin must verify an execution bridge and its input contract for the exact capability ID and revision. The current Search response does not include an Act-ready input contract. Until that bridge is verified, do not start the listing through Act. Search does not authorize work or spending.
Provider capabilities and sources
- Scale AI · Source annotation experts
- Source annotation experts. Confirm the provider agreement and supported execution route.
- Scale AI: expert data
- Surge AI · Provide independent expert review
- Provide independent expert review. Confirm the provider agreement and supported execution route.
- Surge AI: expert workforce
- Factory AI · Build the annotation data pipeline
- Build the annotation data pipeline. Confirm the provider agreement and supported execution route.
- Factory: coding Droids
- Label Studio · Review labels and adjudications
- Review labels and adjudications. Confirm the provider agreement and supported execution route.
- Label Studio: annotation and evaluation
Coordinate annotation, independent review, and the data pipeline
With scope and accounts approved, Muse coordinates Scale AI annotation, Surge AI independent review, Factory AI pipeline work, and Label Studio acceptance through verified Darwin capabilities. Provider-managed experts do the labeling; the AI data lead retains responsibility for the rubric and acceptance decision.
Calibrate the experts before releasing the batch
Muse starts the authorized Scale AI and Surge AI assignments with the sample, rubric version, data-access limits, and acceptance criteria. The providers return their interpretations of ambiguous examples before the team commits the remainder of the work. The buyer records each clarification so the same issue does not produce conflicting instructions in separate conversations.
The calibration sample is part of the 500-case scope, not an additional hidden batch. When instructions change, identify which already-labeled records need rework. An agreement on a revised rubric is not proof that every example now meets it; the subsequent review still has to check the actual labels.
Build an import path the buyer can inspect
Factory AI works in an authorized repository to validate unique IDs, required fields, permitted labels, and export encoding. Its changes go through code review and run against fixture data before production exports are imported into Label Studio. Annotation rationale and provider batch identifiers remain attached to the records.
A failed schema check returns a specific error list to the responsible provider. Code can check that a rationale exists; it cannot establish that the rationale is sound. That judgment belongs to the annotation experts and reviewers, with the buyer deciding how unresolved disagreements affect acceptance.
Sources: Factory: coding Droids · Label Studio: annotation and evaluation
Review independently and resolve disagreements
Surge AI reviews the contracted coverage without silently copying the original label as its own judgment. Label Studio holds the buyer’s comparison and adjudication record, while Scale AI receives actionable correction requests. Muse tracks whether a disagreement was resolved, excluded with a reason, or left open. Production data is not released simply because a provider reports the assignment complete.
From selected capabilities to coordinated work
The calibrated batch, independent review, tested export, and unresolved questions are ready for acceptance.
Darwin Act API
With scope and accounts approved, Muse coordinates Scale AI annotation, Surge AI independent review, Factory AI pipeline work, and Label Studio acceptance through verified Darwin capabilities. Provider-managed experts do the labeling; the AI data lead retains responsibility for the rubric and acceptance decision.
- 1
start_actionStart one scoped assignmentStart the authorized Scale AI assignment using its verified capability ID, revision, agreed scope, and permitted inputs.
- 2
get_actionRead progress and recover the same workRead current progress, review questions, and returned artifacts for each assignment.
- 3
continue_actionSupply a requested handoffReturn the specific approved answer or requested revision while preserving the agreed scope.
- 4
end_actionClose the work and read its outcomeClose accepted work with the current revision after the buyer checks the deliverables.
Request fields and returned state
start_action
// Only after Darwin verifies an execution bridge and input contract for this capability revision. Persist startRequestId; reuse it for retries.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
name: "start_action",
arguments: {
capabilityId: selectedCapability.capabilityId,
capabilityRevision: selectedCapability.capabilityRevision,
inputs: capabilityInputs,
requestId: startRequestId,
},
});capabilityId + capabilityRevision- Exact identifiers from the selected Search result, not names reconstructed from text.
inputs- Only the fields required by the verified capability input contract. The task brief below describes the work, not a universal JSON schema.
requestId- A new idempotency key for this assignment. Reuse it only when retrying this same start.
Retain actionId, revision, lifecycle, status, and availableActions. An accepted or running Action is not completed work.
get_action
// Use the actionId returned by start_action. Read state before deciding on the next operation.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
name: "get_action",
arguments: {
actionId,
},
});actionId- The identifier returned by start_action. Reuse it after an interruption; do not start a duplicate task.
Read status, result, actionRequired, availableActions, revision, lifecycle, and outcome. The result content depends on the selected capability.
continue_action
// Call only when availableActions includes update. Persist updateRequestId and retry only the identical update.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
name: "continue_action",
arguments: {
actionId,
message: handoffMessage,
requestId: updateRequestId,
},
});actionId + message- Send the requested clarification or safe output references to the existing Action, only when availableActions includes update.
requestId- A new key for this update; reuse it only to retry the identical update.
Reread the Action after the update. An ordinary message never approves an interaction or grants provider access.
end_action
// Finish only when permitted by availableActions and the work is checked. Persist endRequestId; read until lifecycle is ended.
// mcp is your connected, authorized MCP client.
await mcp.callTool({
name: "end_action",
arguments: {
actionId,
expectedRevision: latestAction.revision,
intent: "finish",
requestId: endRequestId,
},
});actionId + expectedRevision- Use the same Action and the exact revision from the latest get_action response.
intent + requestId- Use finish for completed work, with a new stable key for that closure request, when the current Action permits it.
An ending lifecycle has no final outcome. Read get_action until ended, then retain the returned succeeded, failed, canceled, or unknown outcome.
Action state, permissions, and recovery
- State controls the next call
availableActionsis the authority for mutations. Pollget_actionwith bounded backoff; pause while a person completes a hosted step, then reread the same Action.- Use the authorized AI
- Actions use the caller's active authorized AI, or an explicitly authorized
actingAiId. Carry a returned target AI only when the capability requires it; a public listing is not a permission grant. - Keep secure steps separate
- Darwin OAuth grants the approved scopes. Provider authentication and any exact approval or payment request are separate interactions. Use the first-party
webLink; never send credentials in task messages. - Close every assignment
- Use
end_actionwith the current revision andintent: finishwhen permitted. Followendingtoendedthroughget_action. A recorded outcome does not replace checking the delivered work.
Accept a usable dataset with a traceable review history
Scale AI supplies annotated examples, Surge AI supplies independent judgments, Factory AI supplies tested import and export code, and Label Studio holds the buyer review record. Muse reconciles those deliverables before the data lead accepts the batch.
Receive a dataset with an inspectable review history
The handover joins each example to its annotation, rationale, rubric version, review status, and any correction. Factory AI’s validation report establishes that the files can be consumed by the intended pipeline. The data lead samples accepted records in Label Studio and checks the review evidence from Scale AI and Surge AI before signing off.
The useful outcome is a batch that can be traced and corrected. A clean export does not demonstrate better model performance, and this engagement does not include a model-training result. The buyer can use the accepted examples in a later training or evaluation run with its own safeguards.
Accept the contracted work against explicit checks
Deliver 500 accepted examples. Track exclusions separately and replace rejected cases within the agreed scope so the accepted count does not quietly shrink. Confirm the independent review coverage, adjudication history, and reproducible schema validation. A batch with missing rationale or unexplained count differences goes back for correction before it is recorded as accepted.
Calculate the net value of coordination time saved
For one 500-example batch, the pilot estimate reduces buyer coordination from 32 to 21 hours. Eleven hours at an assumed $100/hour represents $1,100 of recovered team capacity; subtract a $200 cost allowance for incremental coordination tooling to estimate $900 in net value for the batch. Annotation and review fees remain in the pilot budget in both approaches.
Measure briefing, export reconciliation, and review coordination against the 32-hour baseline. Calculate net time value at $100 per hour after coordination costs, and reconcile expert invoices separately. The estimate values recovered team capacity; it does not reduce the agreed annotation fee.
The result and the effort behind it
500 accepted examples are delivered with annotation rationale, review decisions, schema checks, and a reconciled cost record.
- 1. Search
Find the right capabilities
Comparable provider proposals define qualifications, annotation scope, review coverage, software work, and the full pilot cost.
- 2. Act
Coordinate the handoffs
The calibrated batch, independent review, tested export, and unresolved questions are ready for acceptance.
- 3. Outcome
Return the complete result
500 accepted examples are delivered with annotation rationale, review decisions, schema checks, and a reconciled cost record.
Less searching. Fewer manual handoffs.
Estimated buyer effort for the same scope and acceptance criteria. Track actual hours during the pilot to compare with these estimates.
Time to a reviewed selection
12 working hours without Darwin; 6 with Darwin in the pilot estimate.
Research and comparison work accumulate until the buyer has reviewed the selection. Delivery time is separate.
Human effort across the same scope
32 staff-hours without Darwin; 21 with Darwin in the pilot estimate.
Work roles may belong to the same person or run in parallel. Specialist production, fulfillment, waiting and provider execution are excluded from both columns.
Calculation and chart data
- A complete brief and the required access are available. The same buyer acceptance checks apply in both approaches.
- Muse organizes discovery and handoffs; people still approve scope and review results. Estimated savings come from research and coordination.
- Discovery is included in total effort. Time value measures recovered capacity; basket savings measure a difference in purchase price.
- Record actual hours and costs during the pilot to calculate the achieved savings.
| Milestone | Without Darwin | With Darwin |
|---|---|---|
| Brief | 2h | 2h |
| Research | 7h | 4h |
| Compare | 10h | 5h |
| Select | 12h | 6h |
| Role | Without Darwin | With Darwin |
|---|---|---|
| Data lead | 14h | 9h |
| Data operations | 10h | 6h |
| Privacy review | 4h | 4h |
| Purchasing | 4h | 2h |

