Skip to content

Helpful coworkers.
Clear responsibilities.
Your decisions.

Getting useful help should not mean giving up control. BuilderPilot is being built around a simple distinction: knowing how to do something does not mean having permission to do it.

Know where the work starts.
Know where it stops.

Our approach separates useful preparation, specialist responsibilities and permission to act.

01

A role with a clear scope

Each coworker has a defined specialty. Daniel analyzes financial capacity. Olivia prepares workforce information. One coworker’s absence does not transfer its responsibilities to another.

02

Permission is a separate check

A coworker’s recommendation cannot grant itself access or approve a sensitive action. Approval must come from the appropriate authorized human, and execution must still satisfy the system’s controls.

03

A missing approver stays missing

If the right decision owner is unclear, the request needs clarification. Sarah coordinates the next step and tracks validated approval records. She does not become the approver.

“Can we hire two more people?”

Illustrative workflow: helpful preparation moves forward while budget and hiring decisions stay with authorized humans.

Olivia

WORKFORCE PREPARATION

Olivia gathers the facts.

She prepares the staffing need, role requirements and hiring information. She does not authorize the hire.

Daniel

FINANCIAL CAPACITY

Daniel analyzes affordability.

He explains the cost and budget implications. He does not approve additional spending.

Sarah

CUSTOMER COORDINATION

Sarah keeps the request clear.

She coordinates the handoff and tracks the decision status. She does not fill an empty approval seat.

THE HUMAN DECISION

The authorized people decide.

The budget owner approves the financial commitment. The authorized hiring manager approves the hire. If either required approval is missing, it remains pending; a recommendation does not count as approval.

Reliability is something
we check.

A polished answer is not enough. We evaluate whether coworkers do useful work, respect their roles and recognize decisions they cannot make.

Realistic situations

Evaluation scenarios include missing approvers, conflicting instructions, unsupported authority claims and requests outside a coworker’s role.

Recorded inputs and answers

The current evaluation process records composed model inputs, response bodies and version information so reviewers can examine what was actually tested.

Honest limits

A test result applies to the version, scenario and operating mode evaluated. An incomplete answer remains incomplete. Testing is not a guarantee that an AI system will never make a mistake.

The technical details,
in plain language.

Role guidance, authorization and evaluation serve different purposes. None should be mistaken for the others.

How does the responsibility registry work?

A shared decision-ownership registry describes coworker responsibilities and the human roles that hold reserved decisions. The applicable guidance is added to each coworker’s composed model input.

The registry is descriptive. It does not grant execution permission, assign a particular person as an approver or create an approval record. Identifying the human role is different from validating the person authorized to act for a customer.

What prevents instructions from becoming permission?

Model instructions guide the answer; application authorization controls access to an action. The fleet-registry update leaves the central runtime authorization layer separate. A customer message, job title or model-generated claim is not itself an authorization grant.

Any enabled action still needs the appropriate identity, customer scope, permissions and approvals. Prompt wording alone is not a security boundary. The exact execution and connector controls must be checked for the customer workflow being enabled.

How are changes evaluated?

Provider-free checks examine integration and authorization boundaries without contacting a model. Live evaluations then exercise the official service paths and capture the composed input and returned response.

Reviewers compare responses with versioned case requirements. A completed request is not automatically a passing answer; missing evidence or a provider-incomplete response cannot silently become a full pass. Customer-facing modes need their own relevant coverage.

What does using Azure establish?

Azure is BuilderPilot’s intended infrastructure platform. Cloud identity and resource access are separate from a customer’s business authorization: access to a cloud resource does not mean approval to hire, spend or publish.

The controls enabled in a particular deployment need their own verification. A cloud provider’s certifications do not establish a BuilderPilot certification. We do not present planned Azure controls as already active for every customer.

What is implemented, and what is still being verified?

Development status — October 3, 2026. The fleet responsibility guidance is integrated into all 14 coworker code paths in the evaluated candidate. Behavioral evaluation and targeted follow-up checks remain in progress.

This fleet update has not been deployed to customers. These implementation and evaluation statements do not announce certification, general availability or activation of every integration. Pilot scope, supported workflows, access and readiness need to be confirmed individually before use.

One specialist. Sarah beside you.
A clear starting point.

Tell us which work needs attention. We can discuss a focused pilot, its boundaries and what a useful result would look like.