Innovation & automation leadership · Multi-agent decision support
Innovation Board
A board of six AI agents assesses the commercial potential of a software product idea and prepares a brief for the person who holds the budget. Every claim must quote the evidence, invented figures are struck from the record, and code computes the development estimate, the cash flows, ROI, NPV and payback.
1 · Choose a product idea
2 · Research (simulated)
The research is simulated: the sources are a fixed synthetic evidence pack of eight to ten short documents per idea, shown here as each agent “retrieves” them. Before the debate, code screens the pack and strikes figures from promotional sources that give no method or primary data. A production version would replace the pack with web search, internal data connectors and document retrieval, keeping the same citation check.
Sources appear here when the board convenes.
3 · Board debate
- Press “Convene the board” to watch six agents debate. Each turn is one request; the server checks every quote and figure before the scores move.
4 · Brief for the decision-maker
Assessment profile (0 to 100)
Development time and cost (PERT, about 80% range)
Financial case: three scenarios over three years
Sensitivity: what moves the base-case NPV
Key assumptions (checked against the evidence)
What would change the board’s mind
Cheapest next experiment
Dissent log
Audit trail
Human decision
The board’s classification is advice. You decide. Your choice is shown on this page only: nothing is stored or sent.
How it works
- Orchestrated by code: the page drives a fixed turn order (three opening assessments, an audit with three responses, the financial case in two steps, then the Chair) with a hard cap of ten turns. Each turn is one request to a Cloudflare Pages Function.
- One agent, one compact prompt per turn: the Function builds a short prompt for DeepSeek (OpenAI-compatible API, small output limit, thinking disabled) from the company profile, the evidence pack and the run state. Visitors only pick an idea from a list, so no free text reaches the model.
- No server session: the run state travels with each request, signed with an HMAC. The server validates its structure and signature before acting, so altered claims, estimates or assumptions are refused.
- Guardrails in code, not in the prompt: every quote must be a real substring of the cited source; any figure an agent states must appear in a source that passed the screen or be computed by code in this run, or the statement is struck as an “unverified figure”; estimates and assumptions are range-checked; code computes every score, estimate and financial result and applies the investment policy.
- Never fails because of the LLM: if the key is missing, the provider errors or a turn is invalid twice, the page replays a recorded live run through the same checks and computations, and the badge says “Recorded run”.
Formulas (agents propose, code computes)
- Dimension score (strategic fit, market pull, feasibility, risk) = 20 × mean of the accepted 1-to-5 scores for that dimension, or 0 if none
- Evidence strength = 100 × (accepted ÷ claims made) × Σ q · score ÷ (5 · Σ q), over accepted claims, with source quality q = 1 primary, 0.6 secondary, 0.2 promotional
- Financial return = clamp(40 + 30 × base-case ROI, 0, 100)
- PERT per work package: E = (o + 4m + p) ÷ 6, σ = (p − o) ÷ 6; total σ = √Σσ²; range = E ± 1.28σ (about 80% confidence)
- Duration in weeks = person-days ÷ (squad size × productive days per week); cost = Σ person-days × role day rate
- Customers: C₀ = 0, Cₜ = Cₜ₋₁ × (1 − churn) + adoption; revenue or saving in year t = price × (Cₜ₋₁ + Cₜ) ÷ 2; net cash flow = revenue − running cost
- NPV = −development cost + Σ net cash flowₜ ÷ (1 + discount rate)ᵗ, t = 1 to 3; ROI = (Σ net cash flow − development cost) ÷ development cost
- Payback = the month in which cumulative cash flow, starting at −development cost, reaches zero (linear within the year)
- Scenarios: pessimistic and optimistic take each assumption from the cautious or favourable end of its evidence range and the high or low end of the cost range; the base case uses the Finance Analyst’s verified values and the expected cost. Sensitivity moves one input at a time across its range.
- Tiers, first match wins: Invest now if strategic fit ≥ 60, evidence strength ≥ 60, base NPV > 0 and payback ≤ 24 months; Run a pilot if fit ≥ 50, evidence ≥ 45 and base NPV > 0; Explore further if the optimistic NPV > 0; otherwise Park.
Limits
The figures illustrate a method on invented data and are not forecasts. The company, the ideas, the evidence packs, the rate card and every number are synthetic, and the research phase is simulated. A three-year model with four assumptions leaves out tax, inflation, financing, cannibalisation and the cost of the squad’s time elsewhere. The language model can still misread a source within the rules; the checks only guarantee that what counts is traceable to the evidence or to the code, not that it is right. A person makes every investment decision.
Source code and tests: github.com/nepryoon/innovation-board (opens in a new tab)