Facts Are Hindsight. Roadmaps Are Beliefs.

Managers Are Gamblers and Accountants Are Historians

Robust action under uncertainty is recommendable. Yet, how are facts, beliefs, opinions, expectations and so forth similar and different and how to separate the essential from the unlikely to reach that set of Robust action alternatives and ultimately the plan and action?

When someone in a product or strategy meeting says, “I believe this feature will increase customer demand by 20%,” what are they actually doing?

Depending on who you ask, they are either issuing a data-driven prediction, placing a speculative bet, making an educated guess, or simply manifesting their hopes into the universe.

There is a wide collection of future-facing words. Inside an enterprise, these words mean wildly different things depending on which floor you sit on.

Product planning often includes suggestions like adding a new dashboard widget or other pet feature will increase customer demand by 20%.

The statement sounds precise, authoritative, and completely grounded in strategic foresight. In reality, it is likely to be a personal hunch into the quantitative dialect of executive decision-making.

Tech companies run on these future-facing claims and suggestions, yet rarely acknowledge what they actually are. Depending on who looks at the proposal, the exact same statement qualifies as a user-centric vision, a risky engineering distraction, or a binding revenue commitment. Words like belief, expectation, forecast, and bet get tossed around inter-changeably, blurring the line between empirical certainty and optimistic guesswork.

Before an organization can make sound decisions, it should clarify these words and concepts. Managing product development isn’t about scrubbing away unproven ideas to chase absolute certainty, but about recognizing that every roadmap is inherently built on assumptions.

Understanding how different teams make and process those assumptions is where robust action planning and effective leadership can begin.

The Enterprise Lexicon of Unproven Truths

Every corporate role translates the language of the future through its own functional lens. When a product manager speaks of a belief, they describe a deeply felt product vision anchored in user empathy and qualitative customer feedback. The engineering lead hears that same word as an unsourced hypothesis destined to force scope creep late on a Friday afternoon, while the sales director immediately converts it into next quarter’s baseline quota.

Expectations carry a similar dual meaning across the org chart. Company leadership issues an expectation as a polite euphemism for a non-negotiable directive, whereas middle management absorbs it as a rapidly ticking deadline. When finance builds a forecast, they assemble a complex statistical model across dozens of spreadsheet tabs, while the operational teams treat that exact forecast as corporate astrology wrapped in business casual.

Strategic bets reveal a fundamental divide in risk perception across executive tiers. A chief executive views a bet as a bold, visionary move that demonstrates market leadership to investors and the board. Conversely, a line manager sees that same bet as a high-stakes coin toss where heads guarantees business as usual and tails triggers an organizational restructuring.

Finally, the weight of an opinion depends entirely on who voices it in the room. When the highest-paid person in the conference room states an opinion, the surrounding team often accepts it as settled science. When anyone else offers an opinion without supporting data, colleagues instantly flag it as a risky political stance.

Managing Beliefs Instead of Pretending They Are Facts

Successful product leadership requires managing uncertainty rather than disguising it behind corporate jargon. When teams frame a future feature as a guaranteed fact, they construct rigid development roadmaps around pure illusion. Product managers achieve far better outcomes when they treat every forward-looking statement as an explicit hypothesis.

First, teams must explicitly calculate the financial and operational cost of being wrong before committing resources to a feature. Second, product organizations must design the smallest possible experiment that can validate or invalidate the core assumption in market conditions. Finally, leadership needs to define the exact metrics that will force the team to abandon a failing bet before sunk costs accumulate.

Effective management never eliminates beliefs in favor of rigid historical data. Strong leaders practice deliberate belief management by converting vague optimism into testable bets, using current discovery cycles to manufacture future facts.

Why Fact-Based Management Won’t Save Your Product

Modern product organizations elevate fact-based management to absolute gospel. Executives urge teams to move fast, rely exclusively on telemetry, and purge human intuition from decision-making. While fact-based decision-making sounds disciplined and risk-averse, it masks a fundamental reality: facts exist exclusively in the past.

A fact represents an event that has already occurred and can be and was measured.  Analytics platforms log facts, engineers measure them, and finance teams audit them.

Facts deliver total reliability and zero variance, yet they remain entirely incapable of dictating what a team should build tomorrow. Operating strictly on fact-based management does not produce foresight; it delivers perfect 20/20 hindsight wrapped in a dashboard.

Facts serve as historical bookkeeping entries for accountants, historians, and post-mortem reviews. Conversely, beliefs, expectations, and strategic bets form the core material of future product development. Beliefs can include topics that are hard or impossible to measure – yet.

Product managers operate squarely in that unwritten future, making the deliberate management of unproven beliefs their primary responsibility. Facts are necessary, but they do point to the past – like horseless carriage versus automobile.

Ultimately, navigating the unknown isn’t about eliminating speculation—it’s about managing it with rigor. While historical data tells us where we’ve been, growth requires taking calculated leaps into what comes next. Ainolabs provides tools to help people and organizations structure, track, and evaluate decisions under uncertainty.

Ground future bets in systematic evidence and continuously refine assumptions based on past learning, and you can transform fluffy guesses into a continuous engine for intelligent strategic planning, robust action and execution.

This post was written and drawn with the help of generative AI tools.

Doubt, jagged intelligence and Robust Action

Companies – probably anyone – would like to make informed, robust decisions in the face of uncertainty that result in reasonable outcomes in all or most plausible future scenarios. In other words strategic, flexible measures that accomplish immediate goals while deliberately preserving long-term options.

Jagged intelligence refers to the uneven performance of AI models. They excel in some tasks, but may hinder human effectiveness in others. I saw the phrase in an New York Times article, and read the background study mentioned there.

Human consciousness has been tied with thinking ever since and possibly before the utterance Cogito, ergo sum. With AI systems, some psychologists have paid closer attention to the actual thought process and suggest that the process is at least as important as the outcome of thought. Doubt, mistakes, learning from experience are essential and so the adage should be I doubt, therefore I think, therefore I am.

Jagged intelligence and doubt and thought processes that can be explained are necessary for robust decisions. Decisions that are made while taking into account the limitations of the thought processes will end up more robust when future events fold in unexpected ways.

Ainolabs Belief Management, or Cognitive Fitness tools are built to help collect, capture, articulate and handle doubt and to illuminate peaks and valleys of jagged intelligence. Perhaps they can reveal the occasional black swan waiting in the darkness.

Belief management for people

The previous posts have been about architecture, system design, complexity, computational cost, and other technical and intrinsic matters about building autonomous systems.

This installment focuses on the people, and what beliefs, assumptions, expectations could mean to enterprise employees. Their work, incentives, business success, and company performance under uncertainty is discussed.

The result is an approximate guideline and system design for building Ainolabs Cognitive Fitness tool. The internals is expressed “Belief management”; that probably will be replaced with something less theoretical, e.g “Cognitive Fitness” or “Decision Intelligence” . These should connect to user outcomes instead of epistemology.

Please let me know if this piece or any other is complete non-sense. If it is not – enjoy!

World Model is a set of Beliefs

Beliefs, assumptions, scenarios expectations, and other similar expressions refer to possible futures events with some scope, time, likelihood, and potential outcomes and results.

A company’s World Model should be populated with such information. We call those beliefs, and propose that they are entered in a Belief Manager. This potential system component is outlined in the attached next entry of the Ainolabs architectural components.

World Models in Enterprise Automation

The previous post about autonomous systems discussed World Models in a narrow domain. My goal is to outline, understand and eventually build automated workflows in enterprise environments. The following discussion about Enterprise World Models, Belief Management and Robust Actions v.s. perfect foresight is the next conceptual step

Thinking about World Models

I have spent time in figuring out how the general CMLS architecture would apply to workflows in an enterprise. To get to an understanding I ended doing intellectual – or at least mental – circles and spirals, and needed to start from basics. So here are a few pages of a students journey through enterprise and a bit of technical architecture and design views about how to design an autonomously driving vehicle or how to automate an enterprise workflow.

The text is in an attachment. I don’t like the constraints and defaults of WordPress for slightly more serious work.

Smart system cost

We will create a somewhat concrete, yet still hypothetical example.

We’ll illustrate the total cost over a year (365 days) of operation for the following scenario:

The system is a free-roaming four legged front-loader robot.
It has two arms to handle parcels and packages. The arms have sensors for weight and surface characteristics – “gentle touch” not to crush materials or human operators.

The robot operates inside a facility, warehouse or factory. It has a visual, lidar and radar system for observing and navigating the environment. That is subsystem number 1, “Navigation”.

The front-loader has a system to handle propulsion, the four legs. That is subsystem 2, “propulsion”.

The front-loader has a system to handle the materials in the warehouse with the two arms. This is subsystem 3, “payload processing”.

The front-loader has a separate visual, audio and textual UI system for interacting with human workers in the facility or elsewhere (remote connection). The front loader’s UX is friendly and based on state of the art Human Computer Interface practices. This is subsystem 4, human interaction.

The front loader is re-trained every 24 hours, i.e. 365 times per year. The initial training material for subsystems 1-4 is almost completely disjoint. Each separate training material has a high signal to noise ratio. The system is expected to handle 20 human interactions and 10 parcel operations per hour.

With these characteristics, we will compare between all ML eggs in one model v.s. four disjoint models.

Architecture alternatives

A. Monolithic model
  • One large multimodal model handling all subsystems jointly.
  • Shared latent space across navigation, control, manipulation, and HCI.
  • Retrained end-to-end every 24 hours.
B. Multimodel system
  • Four specialist models:
    • M1M_1​: Navigation
    • M2M_2: Propulsion
    • M3M_3​: Payload processing
    • M4M_4: Human interaction
  • Lightweight integration layer:
    • Task router
    • Shared state abstraction
  • Each subsystem retrained independently every 24 hours.

Because training data is almost completely disjoint and each subset has a high signal to noise -ratio, this is a best-case scenario for modularization

Parameter and scaling assumptions

These are deliberately conservative and internally consistent.

Model sizes

Let:

  • Monolithic model size: Pmono=1010  parametersP_{\text{mono}} = 10^{10} \;\text{parameters}
  • Each specialist (thanks to disjoint, high-SR data): Pi=1.5×109P_i = 1.5\times10^9

Total specialist parameters:Pi=6×109\sum P_i = 6\times10^9

Modular storage is smaller, not larger. That is realistic in this case since domains barely overlap.


Training cost scaling

We’ll assume that training cost is proportional to the number of parameters P in a model. CtrainTPC_{\text{train}} \propto T \cdot P

Let one full training of the monolith cost:Ctrain,mono=1.0  cost unitC_{\text{train,mono}} = 1.0 \;\text{cost unit}

Then per full retraining:Ctrain,multi=0.15  per subsystem0.6  total per dayC_{\text{train,multi}} = 0.15 \;\text{per subsystem} \Rightarrow 0.6 \;\text{total per day}

This reflects:

  • smaller models,
  • higher SR,
  • no cross-domain entanglement.

Inference activity volume and cost

Per robot:

  • Human interactions:
    20×24×365=175,20020 \times 24 \times 365 = 175{,}20020×24×365=175,200 / year
  • Parcel ops:
    10×24×365=87,60010 \times 24 \times 365 = 87{,}60010×24×365=87,600 / year

Assume each event requires:

  • Monolith: full model inference
  • Modular: 1–2 specialists activated, average = 1.5

Assume inference cost ∝ active parameters.

  • Monolith inference cost per event: Cinf,mono1010C_{\text{inf,mono}} \propto 10^{10}
  • Modular inference cost per event: Cinf,multi1.5×1.5×109=2.25×109C_{\text{inf,multi}} \propto 1.5 \times 1.5\times10^9 = 2.25\times10^9

That is ~4.4× cheaper per interaction for the Smart system based on multiple integrated models, or Docker for AI.

Almost there: Five-year Total Cost

Training cost (5 years)

ArchitectureDaily costDays5-year total
Monolithic1.018251825
Multimodel0.618251095

Training savings: ~40%


Inference cost (5 years)

Total interactions per year:175,200+87,600=262,800175{,}200 + 87{,}600 = 262{,}800175,200+87,600=262,800

Five years:1.314×106  events1.314 \times 10^6 \;\text{events}1.314×106events

ArchitectureCost per event5-year total
Monolithic1.01,314,000
Multimodel0.225295,650

Inference savings: ~4.4×


Storage & integration (5 years)
ComponentMonolithicMultimodel
Model storageHigh (10B params)Moderate (6B params)
Integration infraMinimalModerate
Net effectBaseline+5–10% overhead

We will conservatively add 100 cost units to multimodel TCO.

Final TCO comparison (5 years)

Cost componentMonolithicMultimodel
Training1,8251,095
Inference1,314,000295,650
Storage + integration~0+100
Total TCO~1,315,825~296,845

And conclusions:

So what did we do and say here?

We outlined a theoretical, yet plausible system, and compare two alternative ways to build that. The architectures we compared are a single large model that handles everything (monolith), and a system built of components, i.e small independent ML models that are integrated (multimodel architecture).

The multimodel architecture is ~4.4× cheaper over 5 years, dominated by inference cost savings.

Why modular wins decisively here

  1. Disjoint, high-signal domains
    No representational duplication penalty.
  2. Daily retraining
    Training efficiency compounds strongly over time.
  3. Sparse activation at inference
    Only the relevant subsystem runs per task.
  4. Embodied system
    Most tasks are local (navigate, lift, talk), not global reasoning.

This is almost the ideal use case for modular intelligence.

In this scenario, a monolith is paying a tax for generality it does not use most of the time.

Development cost of an AI system

Ainolabs talks about systems, entities that are operated for defined purposes and with expectations of generating value to their developers, owners, users and other stakeholders. Value depends on the business case. The cost can be estimated, and we’ll get to this formulaic expression below:

Systems are developed, launched, updated, supported and eventually replaced or discarded. They have a life span, and with life span there is the life time cost. We try to estimate that on a somewhat abstract level of complexity, as discussed in Computer Science.

There are two ways to build an AI system. One is the current, train everything in a single model, possibly create agent instances in that, package the model and release and run.

It is also possible to build a system of interconnected smaller models, minimodels. That is cost effective when the system functionality splits into disjoint subdomains, and especially if the combined dataset would have very low entropy, or, put in signal processing terms, signal to noise ratio (SR) is very low in the aggregate data set, but high in the parts when compartmentalized.

Building a system with many models incurs general overhead. In the Ainolabs architecture that means the burden of Content Classifier. This adds to the processing of each input, and should be balanced by the reduced cost of processing in the separate models (Learning Processes).

Scientish-like complexity estimates of development

By AI system we mean a system that uses Machine Learning technology as an essential element. A system could be built, as mentioned above as a monolith, or as a collection of separate smaller models. The development of an ML (AI) system is mostly related to training, and so the following estimates the order of magnitude costs of that.

We’ll use the following terms:

  • T = number of training examples
  • d = feature dimension per example
  • P = number of model parameters (often ≈ scale of the model)
  • E = training epochs / outer iterations (or rounds/trees, etc.)
  • k = clusters or neighbors (when relevant)
  • F = per-example forward+backward FLOPs (architecture-dependent; ≈ proportional
    to P for dense nets)

In machine learning, P represents the total number of trainable parameters — the variables
that the training process adjusts to minimize loss.

Taking noise into account

The number of parameters in the model, P, depends on the quality of the data. This has an impact on the model, training and semantic model capacity. Data with a lot of noise and no signal results in a very easy model, but with no practical value – just consider the ultimate white noise, cosmic background 4K radiation. P is conceptually a measure of information in a model.

If the training data as an aggregate is noisy – e.g. different sensory systems for different purposes look like noise when combined – splitting it might help.

Next we try to create a way to express and predict P in terms of dataset characteristics like size T , signal-to-noise ratio (SR), and desired accuracy A.

For that we need to define

For a model the number of parameters P needed is roughly proportional to the amount of true signal in your data and inversely proportional to the noise and desired accuracy.

We’ll skip a few steps and claim that the number of parameters is proportional (“Oh”-notation) to the size of the corpus T, signal to noise ration SR, and desired accuracy A. Exponents depend on the specific situation:

And with that it is possible to express the order of magnitude of training a model as

Divide et impera, or split the system into parts

The full dataset T might be large, could consist of orthogonally different data (audio, video, and temperature, for example). In such a case it would make sense to split the task – LP1, LP2 etc in the architecture.

Next total dataset has size T, and it is divided it into disjoint subsets:

Training cost for comparison between a monolith and multimodel system are

With these two it is possible compare the effectiveness of different or combined data sets for an operational system that utilizes Machine Learning technology – an AI system.

Elements of Trust


An essential part of the architecture for continuous machine learning are the trust relationships between components and participants of a system and with possible external parties – contributing actors.

Content credentials https://contentcredentials.org/ attempts to address trust by placing a mark on content. This assumes that the recipient trusts the mark, and that the marks are not abused by malevolent actors. Both are quite possible.

A system that relies on any information or action by other parties needs to determine and maintain trust information about others. Trust is a multidimensional entity, including trust in someone’s veracity, competence, good will, transient factors, or other attributes. A human example is trust based on group membership.

The architecture contains an element “Trust level adjustment”. That refers to a simple process outlined in the illustration below:

This reflects trust relationships between two agents (components), i and 0 (self). The relationship is asymmetrical on every attribute, attributes between agents may differ, and the trust levels are a function of past experience.

With such mechanism it is possible to build systems that gradually learn appropriate trust levels with their environment, The correct trust level is essential for continuous operation. It is argued that without that capability it is not possible to build an artificial intelligence system that can operate in an open environment – i.e. so-called AGI can not work without a proper trust adjustment technology in place.

Architecture for Continuous Machine Learning

Ainolabs architecture for Continous Machine Learning Systems (CMLS) is designed to support non-interrupted operation, modular construction, system and data integrity, scalability in system design and operation and robust continuous operation.

These capabilities result in a machine learning system that can meet requirements on privacy, confidentiality and adapt to changing circumstances.

The blocks and the whole CMLS entity operate as a collection of re-enforcement learning engines. The individual inputs and outputs together with the memory blocks form a learning loop over time.

A CMLS receives inputs I0  … (to In, for discussion where an upper limit is needed). The CMLS responds with outputs O0 … On, again an endless sequence.  Inputs and outputs match each other so that each Ii results on corresponding output Oi The system can produce an empty output.

The system uses the available memory blocks Short Term Memory and Long Term Memory to keep a log of inputs and outputs. This log is stored as memory imprints in the NN of memory blocks for this discussion. An implementation may also store a traditional transaction log; however that is not relevant to ML discussion.

The system uses each new Input Ik as an cumulative observation (feedback) of earlier input and output sequences I0 .. Ik-I and O0 … Ok-1. This provides a CMLS with an increasing set of cumulative experience that can be used to re-enforce earlier memories or resolve ambiguities and possible contradictions.

Input and Output sequences from 0 to k are marked as I(k) and O(k) in the rest of the text. These denote the ordered sets containing items 0 to k.

A transaction is defined as an Input – Output pair, i.e. Transaction = { I, O) and specifically Ti = { Ii, Oi }. The set of transactions 0 … i is marked as T[i].

CMLS Blocks explained

CMLS Blocks accept Inputs, process those using Output drafts and finally after minimizing the overall Cost function send out the Output.

Output can be empty if the Input is not understood or the system decides it is appropriate to “no comments”.

Stimuli  / input

Input is received via some digital channel. Input can represent anything, including but not limited to electromagnetic radiation, sound in any medium, video (images, separate from light), text, illustrations, diagrams or other information or signal occurring naturally or generated by humans.

Input is associated with information about its source. Source can be unknown. See Source Classification.

Content Classification

Content Classification attempts to label the content into classes stored in Long Term Memory. This classification is context dependent and is affected by recent stimuli. The specific effect depends on implementation and can be e.g. increasing or decreasing enforcement.

Content Classification can impact the memory imprint of Input both in Short Term Memory and Long Term Memory.

Source Classification

Source Classification attempts to label the Source of the Input. It uses information stored in Long Term Memory. Source Classification is context dependent. 

Content Classification can impact the memory imprint of Input both in Short Term Memory and Long Term Memory.

Short term memory

Short Term Memory is keeps track of immediate input – output sequences and other transient pieces of information passing through the CMLS. It should be noted that unlike humans, a digital CMLS Short Term Memory is limited only by available computing power and data storage.

Short Term Memory overflows strong memory imprints to Long Term Memory when it is either full or the Input is a wide match with the Short Term Memory’s Neural Network (substantial part of STM’s neurons fire up). 

Long term memory

Long Term Memory is the mass storage of CMLS. All permanent memory imprints are stored there.  The capacity of Long Term Memory is essential for CMLS – the more synapses, the more knowledge.

Long Term Memory is essential for storage (learn), retrieval of information, and also for graceful forgetting. As CMLS is a digital device, there is no biological degradation. Graceful forgetting is intentional, and can be based on time, source, or other factors. Notably, Long Term Memory stores timestamps with new information and can fade or completely remove memories that are valid only for a limited period of time.

Short Term Memory has a substantial role in time-related graceful forgetting. Fleeting information is kept only in short term memory and will not leave an imprint on Long Term Memory at all. As an example, consider weather or traffic conditions day before yesterday. Those can be recalled for a few days, but now e.g. a month ago unless there was a specific and pressing reason to store the memories in Long Term Memory.

Learning process for new inputs

Learning process takes Input and associated Content and Source classifications and runs learning processes on Short Term Memory. See also Re-enforcement process.

Re-enforcement process for familiar inputs

The Classified Input could be familiar from the past. RE-enforcement checks that by retrieval from Long and Short Term memory. In case the Input is familiar, the existing memory, including metadata about Source, Time and other possibly relevant factors is re-enforced together with the current Input via a learning cycle to either or both memory blocks.

Known entities

Known Entities  block relies on memory blocks to check if the current Source is known from the past. Known Entities will run a learning cycle on Short Term memory about new sources, and both memory blocks for existing entities.

See also Trust level adjustment.

Trust level adjustment

Trust level adjustment compares the current Input is compared with existing memories. If there are discrepancies between current Source and other Entities, and/or between the content. Trust levels to either current Source and/or familiar Entities are adjusted. Adjustment happens through a machine learning cycle on Short and Long Term Memory.

Trust adjustment is unique in impacting Long Term Memory direct in case of changing the trust on a Known Entity.

Contradiction processing

Contradiction Processing is activated to check if the Input and/or current Output draft is contradicting Long or Short Term Memory. In layman terms, Contradiction Processing is checking for disagreements, untruths, lies and differing belief structures for both content and source.

Contradiction Processing may impact Known Entities, Short Term Memory, Long Term Memory and the current Output draft. Contradiction Processing can trigger Trust Level Adjustment.

Contradiction Processing for Transaction Ti considers the full history T[i-1]. The NN is calibrated to minimize the possible difference between the new Input Ii and the tentative Output Oi with known history T[i-1]. Specific implementations can weigh information according to its age differently, which results in different degrees of graceful forgetting and learning new information.

Recipient  Classification

The tentative Output may be fine tuned according to the expected Recipient (the party providing the Input). This may not always be possible direct, but system can make a difference between an anonymous party and Known entities.

Recipient is classified similar to the Source via Known Entities, and Trust Level Adjustment. Notably, if memories in Long Term Memory are labelled confidential to certain Entities, any Output candidate to a Recipient in the same Trust circle would be treated differently than Recipients not in the same trust circle.

Response /  output

The CMLS is an Input – Output system with memory. Accordingly, each Input is matched with an Output. The Output can be empty if the CMLS determines it can’t provide a reasonable answer due to  e.g. confidentiality or plain ignorance.

Output is generated gradually between the blocks. It is determined to be ready once the cost functions is acceptably low, or after a timeout. The timeout value can depend on the nature of the Input, but ultimately there is an upper limit to the time CMLS can take to respond to an Input.

Even if the Output is empty, the system may have generated memory imprints of the Input. Over time this is expected to result in learning and CMLS ability to eventually answer.

Conclusion

An architecture for building scalable systems from modules for machine learning systems has been presented. This is similar to the invention of a subroutine: It is not necessary to build a whole system into one model. Instead, working systems can be composed of Domain Specific Models that are set to interact in an efficient and suitable combination.