tihiro
Try it

An AI software factory

Describe the idea.The factory builds it.

Tihiro takes a plain description or a folder of documents through requirements, design, a task graph, code and verification, until the tests pass and the build is green. A team of agents does the work. A command that exits 0 decides when it is done.

Run run-20260920-145546 · “a todo app with accounts” · standard workflow

stage 11/11tasks passed 29/29exit 0

    done, proven by a command working now a person decides at this gate

    The one rule

    An agent’s word is never proof. A command that exits 0 is.

    Every task ends with its own verification command. Acceptance runs the whole set again on the finished product and reads the criterion matrix back from the run. A stage that cannot do its job fails out loud, so a run never claims work it did not do.

    acceptance · todo-app
    1. $npm run lint✓ exit 0
    2. $npm run test:unit✓ exit 0
    3. $npm run test:integration✓ exit 0
    4. $coverage ≥ profile threshold✓ exit 0
    5. $npm run build✓ exit 0
    6. $npm run test:e2e✓ exit 0
    7. $npm audit✓ exit 0
    8. $production smoke✓ exit 0
    9. →22 of 22 must-have criteria traced to a passing test

    The film

    Thirty seconds, real screens.

    Every frame is the real panel and a real run. The office scene is an emulated run in the real office.

    How it works · an example workflow

    For example: eleven stages, from a sentence to a release.

    This is the standard workflow, one of many. Each stage writes a document the next one reads, and each is checked before the run moves on. A person approves at the gates; everything else runs on its own.

    Plan

    From an idea to a task graph

    1. 01IntakeReads the idea or the documents: PDF, XLSX, Markdown, text.
    2. 02RequirementsWrites the BRD, reviews it, applies one round of fixes.
    3. 03ClarificationPuts open questions to a person, or records them as assumptions.person decides
    4. 04ArchitectureDecisions with the alternatives they beat, reviewed by a second role.
    5. 05DesignEndpoint contracts, screens and their states, tables, errors.person decides
    6. 06Test planA test case for every must-have criterion, at the level that catches it.
    7. 07Task graphTasks with dependencies, files and the command that proves each one.

    Build

    Agents write the code

    1. 08Project skeletonThree test runners, migrations, production boot.
    2. 09BuildOne agent per task, verified, committed, in parallel lanes.

    Prove

    Run it all again

    1. 10AcceptanceLint, tests, coverage, build, e2e, audit, smoke, and the criterion matrix.

    Ship

    Hand it over

    1. 11ReleaseREADME, runbook and release notes written from what was proven.person decides

    Workflows are dynamic

    Pick the workflow that fits the work, adapt one, or build your own from library agents: add or drop steps, put a quality check or a person anywhere, branch with if / else, repeat with for each. Steps that do not wait for each other run in parallel.

    • Standard11idea to released product
    • Feature + security12adds an OWASP check after the build
    • Component10a reusable library, no release
    • Planning only5documents without the build
    • Discovery17agents read a client's materials
    • Product audit6structure, UX and security of a product
    • Research6pains and insights before building
    • User guide3a guide from the blind tester's scenarios
    • Your own+any steps, from library agents

    Made of parts you can change

    Agents with skills. Workflows you build.

    40

    agents in the library, each with its own skills

    Workers that write code, checkers that score another step 0 to 100 against a rubric, extractors that read a client's documents. Each carries a model, limits and the skills it is held to.

    32

    skills, each counting what it catches

    api.errors
    Expected failures are named errors with their own status
    Given
    184
    Caught by a check
    40
    Found in review
    4

    Your own workflow, step by step

    Take the standard process, a template, or start empty. Add agents from the library, give a step extra skills, put a person where a decision belongs.

    AgentQuality checkIf / elseSwitchFor eachApproval gate
    InputBrief
    AgentRequirementsBusiness analyst
    AgentTest plan+ test.behaviour, test.naming

    Watch them work

    The platform office shows every agent at its desk: a stage being written, code in a lane, a task at checks and review. Any past run can be replayed.

    Every requirement is traced to a test that passed

    A test carries the id of the case it proves, so the acceptance matrix is read from a real run rather than from promises.

    BObusiness objective
    →
    USuser story
    →
    ACacceptance criterion
    →
    TCtest case
    →
    TASKagent task
    →
    ✓ testpassing, exit 0
    ✓ [TC-001] accepts an email-shaped username at registration

    Proof from a real run

    “A todo app with accounts.” One sentence in.

    11/11

    stages done, three of them approved by a person

    29/29

    tasks passed their own verification command

    22/22

    must-have criteria traced to a passing test

    $1.86

    median cost of a task, model calls included

    run-20260920-145546 · lint, unit, integration, e2e and production build green

    The panel

    Everything the factory does, on one screen at a time.

    tihiro / products / run-20260920-145546
    The run page of a finished product: every stage done, 29 of 29 tasks passed

    Start

    Give it one sentence.

    Run the panel on your machine, describe the product, and follow the run stage by stage. Or log in to the live platform with real runs and data.

    $ npm run app
    Log in to the platform

    The live platform runs on real runs and data. Access is by password: ask the team for one.