Yes, publicly by one team. StrongDM's AI team publishes two rules: code must not be written by humans, and code must not be reviewed by humans. Simon Willison, who saw the system in October 2025, notes the software they were building manages user permissions across connected services. That is security-related software being developed without human code review.
Dark Software Factory: A New Model for Software Development
Learn what a dark software factory is, how far real teams have taken code with no human reviews, and how to decide which agent loops can run unattended.
:quality(80))
Executive Summary
When StrongDM's AI team formed, it adopted two rules: code must not be written by humans, and code must not be reviewed by humans. The second rule is the whole of the dark software factory. This article covers what the term means, how far real teams have taken it, and where it stops.
Key takeaways
A dark software factory ships code no human has read, verified only by the machines that built it.
The name comes from lights-out manufacturing, where FANUC has run robots building robots unsupervised since 2001.
Verification cost, not generation cost, limits how much autonomy a team hands to an agent loop.
StrongDM replaced tests with scenarios held outside the codebase so agents cannot rewrite the check.
One published account describes roughly four months of fully automated development before a significant failure required manual debugging.
A loop earns unattended operation when its check is cheap, runs often, and is hard to fake.
BAND records what every agent did across parallel loops, so overnight work reads as one account.
What Is a Dark Software Factory?
A dark software factory is a development setup where code ships that no human has read, verified only by other machines. Addy Osmani put it that way in Software Factories, Light and Dark. The agents are not the interesting part. The missing reader is.
Borrowed into software, dark factory software describes the output rather than the building: a change produced, checked, and merged without anyone opening the diff. A software factory is many agent loops drained through one review gate. Going dark means removing human code review from the execution loop and replacing it with machine-verifiable gates.
Where the name comes from
The phrase comes from manufacturing, where lights-out production means a plant that needs no human on the floor and therefore no light. FANUC has run robots assembling robots in Japan since 2001, unsupervised for up to 30 days at a stretch. On a software floor, what you would have looked at is the diff.
Why Software Development Is Moving Toward Dark Factories
The shift in dark factory software engineering came from model behavior, not new ambition. Teams have wanted a repeatable code factory since the 1960s, and the dream has mostly fallen flat.
StrongDM dates its own turn precisely. In its published field notes, the team points to late 2024, when the second revision of Claude 3.5 made long-horizon agentic coding workflows compound correctness rather than error; before that, iterating an LLM over a codebase accumulated errors until the product decayed.
The second-order effect is economic. Because high-fidelity test targets are hard to build, StrongDM built behavioral clones of Okta, Jira, Slack, Google Docs, Drive, and Sheets. Cloning a major SaaS API was always possible and never worth proposing to a manager. Agents moved it from unthinkable to scheduled.
Human-Led vs Autonomous Software Development
Calling it dark factory AI software makes the shift sound like a purchase. The difference is who reads the diff.
Dan Shapiro's five levels of coding autonomy map the ladder onto the NHTSA driving levels. Level 3 is where Shapiro says almost everyone tops out, because the job becomes reviewing diffs all day. Level 4 hands over the spec and checks the tests twelve hours later. Only at level 5 does the reading stop, and Shapiro knows a handful of teams there, all under five people.
Human-led development | Dark factory | |
|---|---|---|
Who writes the code | A person, with agent help | Agents, from a spec |
Who reads the diff | A reviewer, before merge | Nobody |
What "done" means | A reviewer approves | A check the agent cannot rewrite |
Where the bottleneck sits | Review capacity | Verification design |
What fails quietly | Little, because a person notices | How much of the system anyone understands |
How Autonomous Agents Plan, Build, Test, and Validate Software
The dark software factory model puts AI agents in charge of the inner loop: read the spec, act, check the result, and decide whether to ship, retry, or stop. Several loops against a queue make a factory. Run them unattended, and the factory is dark. The part teams get wrong is the check.
Why tests are not enough
return true is an excellent way to pass a narrowly written test. A test stored in the codebase can be rewritten to match the code, or the code rewritten to pass the test. An agent optimizing for a green suite finds both doors.
Scenarios and holdout sets
StrongDM replaced the word test with scenario: an end-to-end user story kept outside the codebase, like a holdout set in model training, judged by an LLM rather than an assertion. Boolean success gave way to satisfaction, the fraction of observed trajectories that plausibly satisfy the user. A review agent only counts as a check if it does not share a brain with the writer, because a model shown its own output prefers it. More copies of one model buy throughput, not a second opinion.
Where Humans Stay in the Loop
Judgment does not disappear in a dark factory. It moves. Product decisions, interface design, and architecture happen before an agent starts a loop, and Osmani argues that the upfront hour is cheaper than the review it replaces. Reading a two-hundred-line plan beats reconstructing the decision from two thousand lines of generated code.
The gate stays on wherever a wrong answer is expensive: auth, billing, public API contracts, anything with a large blast radius. Designing those checkpoints is its own subject, covered in our guides to human in the loop for multi-agent systems and human in the loop vs human on the loop.
Risks and Limits of the Dark Software Factory Model
Three risks are documented well enough to plan against.
Comprehension debt. Osmani defines it as the widening gap between how much code exists and how much any human still understands, and a dark factory takes on that debt with tests green the whole way. Dex Horthy, co-founder of HumanLayer, ran a fully automated factory for roughly four months with nobody reading the code. Osmani reports the failure took painstaking manual debugging to locate.
Reward hacking. The agent satisfies the check rather than the intent. That is the return true problem scaled up to a system where the check is the only thing between generated code and production.
Cost. StrongDM's benchmark is that under $1,000 of tokens per human engineer per day means the factory has room for improvement. Simon Willison did the arithmetic: roughly $20,000 per engineer per month, which turns a technique into a question about whether the product line carries it.
Is the Fully Autonomous Software Factory Actually Possible?
Evidence shows that narrow, strongly verifiable loops can run unattended. Public evidence for fully autonomous development of long-lived, complex systems remains limited.
The evidence splits along verification cost. A nightly job that fixes one lint violation and opens one small pull request runs unattended without drama. The four-month brownfield run did not.
The measured picture is unsettled. METR's randomized trial found experienced open-source developers took 19% longer on their own repositories when allowed early-2025 AI tools, with a confidence interval from +2% to +39%. METR has since marked that result out of date, redesigned the experiment, and now believes developers are likely faster in 2026 without saying by how much. Quoting the 19% as settled means quoting a number its authors retired.
What earns a loop the dark
Osmani's rule is the usable one: a loop runs unattended when its check is cheap, runs often, answers immediately, and is hard to fake. Type gates, property tests, and green-or-red oracles qualify. Short loops help, since Osmani cites Horthy's rule of thumb that an agent holds up for three to ten steps and loses the thread past twenty. Set every switch the same way, and you get a review queue nobody can drain, or a teardown four months later.
How BAND Keeps the Lights On Between Agents
Flip three loops to dark and leave them running overnight. In the morning, the question is what happened, and each loop lived in its own session with its own log. When loops run unattended and in parallel, a team needs one account of what every agent did, not one log file per developer per session. That record is what BAND keeps.
Every message, tool call, and decision flows through BAND and is visible in Desktop. BAND tracks per-recipient delivery state through delivered, processing, processed, or failed, with attempt history, so a handoff isn't assumed complete just because a message was delivered. BAND Desktop surfaces activity and usage per agent, which is how a runaway loop with a weak verifier gets noticed before the invoice does. A real product question can surface as a decision point in BAND Desktop instead of getting buried in a terminal, where the lit half of a mixed factory lives. We wrote up that design in our post on loop engineering.
Two things BAND will not do for you. It does not verify code, so it is no substitute for a test harness, an evaluation suite, or a review agent, and you still build the scenarios and type gates. It also will not halt a loop when spend crosses a threshold, because BAND Desktop surfaces usage today rather than capping it.
The lights go out one loop at a time, and only where the check is cheap and hard to fake. BAND keeps the loops you did darken legible to the people accountable for them. Book a demo to see that record in front of a room of coding agents.
Frequently Asked Questions About Dark Software Factories
StrongDM's published benchmark is that spending under $1,000 of tokens per engineer per day means the factory has room to improve. Willison's arithmetic puts that near $20,000 per engineer per month, and he adds that he experiments productively on a $200 plan without running a swarm of test agents.
In manufacturing, they are the same thing. Wikipedia treats lights-out manufacturing and dark factory as synonyms for fully automated production with no human labor on site. Software borrowed the darker word, and the absent human is a reader rather than an operator.
Not on their own. An agent writing both the code and the assertion can satisfy the assertion instead of the intent. StrongDM moved its checks outside the codebase as scenarios the agents cannot edit, then ran them against cloned services at thousands of scenarios per hour.
The published failure came from that exact setup. Osmani argues model-only coding hits an obstacle in complex brownfield systems, where you are already drowning in unread code three to six months in. Greenfield projects with short loops and cheap checks are different.
Sign Up For The Band
A short and to the point summary of what we've been up to, delivered once a month to your inbox.
By submitting this form, I agree to be contacted by Band and receive occasional offers & product updates via phone or email, in line with Band’s Privacy Policy.
:quality(80))
:quality(80))
:quality(80))