A practical system for private employee memory, grounded company knowledge, specialised agents, flexible models, and human approval where it matters.
Why Intricare built this for itself
We had several AI subscriptions across the team. We still spent too much time doing the same work around the AI.
Before a useful answer could come out, someone usually had to explain the product again, paste relevant context, correct the tone, and check whether the model had quietly invented a feature or mixed one product with another. A new subscription gave us another capable model, but it also gave us another empty memory.
That was frustrating in ordinary work. It was worse in work that carried real risk.
A generic reply to a customer can lose trust. An incorrect product claim can create a support problem. A good sales email that ignores the buyer's context is still a missed opportunity. We could not give an assistant access to sensitive company information and then hope it would remember the right things, forget the right things, and stop at the right moment.
The models were not the issue. The missing layer around them was.
We needed a system that could remember an employee's working preferences without treating them as company policy. It needed to answer from approved company knowledge without turning every private conversation into a shared fact. It needed to work across models and providers. And when an agent prepared a customer reply, a financial action, or an operational task, it needed to leave the final decision with a person.
So we stopped asking which subscription to buy next. We built the operating system we wished existed for our own team first.
This is the story of what we chose, why we chose it, and what changed once AI was treated as part of our working system rather than another tab in the browser.
What an AI operating system for business actually is
An AI operating system is the layer that connects models to the people, knowledge, tools, and approval rules of a company. It determines which agent handles a task, which private context it may use, which approved company sources ground the answer, and which actions need human approval.
It is not a replacement for OpenAI, Anthropic, Google, or self-hosted models. Those are model options inside the system. The operating layer keeps memory, knowledge, workflows, and permissions from being trapped inside one provider account.
For Intricare, the system combines Hermes, Multica, Open WebUI, Honcho, and Hindsight. Each component has a different responsibility, which we explain below.
Why we did not simply connect one provider API
Our first instinct could have been to choose one provider and connect its API directly. That would have been simpler, but it would not have solved the problems we were seeing.
A direct OpenAI or Anthropic integration gives a company access to excellent models. It does not automatically provide private employee memory, approved brand knowledge, specialist agents, task tracking, source boundaries, or approval workflows. It also leaves the organisation closely tied to that provider's model catalogue, context handling, account structure, and memory layer.
We wanted the model to be replaceable while the useful company system stayed in place.
That meant putting an operating layer above the provider API. The operating layer would decide which profile is working, which knowledge it can use, what personal context belongs to the user, which tools it can call, and what requires human approval.
We also wanted a model strategy that could grow in three directions: direct provider APIs, a managed multi-model API, and self-hosted models. A direct provider connection is still a valid option for some workloads. It is simply one part of the architecture, not the architecture itself.
Why we selected each part
We did not select these products because a diagram looked impressive. Each one solved a problem we had already experienced.
Hermes: the self-improving agent layer
We chose Hermes because it is more than a model wrapper. It gives us profiles, tools, skills, scheduled automations, integrations, session search, and a learning loop. Hermes can create reusable skills from complex work, improve those skills through use, and preserve useful operational knowledge across sessions.
That self-improving capability matters for an internal system. When we solve a difficult workflow, we do not want the solution to disappear into one chat. We want the agent to turn the lesson into a reusable procedure that future runs can use and improve.
Hermes also gives us provider freedom. We can switch models without rebuilding every profile, tool, skill, and workflow around a new API. That is the main reason Hermes sits at the centre of our design.
Multica: turning agent output into owned work
Hermes can do work, but a company also needs ownership and status. Multica gives us the task and project layer for approvals, assignments, comments, and agent execution.
Without this layer, useful agent work stays inside chat. With it, a customer request can become a draft, an assigned approval task, and a traceable outcome.
Open WebUI: making the system usable by employees
A capable backend is not enough if employees need to understand the infrastructure behind it. Open WebUI gives the team a familiar conversational interface for approved models and agents.
We wanted the complexity to stay in the architecture, not in the daily employee experience.
Honcho: personal memory that can reason over time
We chose Honcho for private, user-specific memory. Its memory model is not limited to retrieving old text chunks. Honcho stores conversations and events, reasons over them in the background, and builds changing representations of people, agents, groups, projects, and ideas.
That background reasoning is the part we think of as Honcho's dreaming or reflection cycle. It can turn repeated interactions into useful conclusions about a person's preferences and working context, which the agent can query later.
This gives each user continuity without forcing the company knowledge base to absorb personal conversations. It also means that changing model providers does not have to erase the employee's useful context.
Hindsight: approved company memory and grounding
We chose Hindsight for the other side of the memory problem: durable, approved company knowledge. Hindsight lets agents retrieve conclusions and source-aware information from a product or company knowledge bank instead of relying on whatever happens to be in the current prompt.
We keep Hindsight banks separated by brand and purpose. That makes it possible to ground a LeadCRM answer in LeadCRM sources without allowing a SalesStack or LinkedFusion claim to slip into the response.
Honcho and Hindsight are deliberately different. Honcho answers, "What does this person tend to prefer or remember?" Hindsight answers, "What has the company approved as knowledge?"
The combination: a system rather than a collection of tools
The value comes from the relationship between these components:
Hermes runs specialised agents and improves procedures
Multica owns tasks, approvals, and execution status
Open WebUI gives employees a simple chat surface
Honcho remembers private user context over time
Hindsight grounds work in approved company knowledge
We could use each product separately. Together, they give us the control and continuity that individual model subscriptions did not provide.
The problem was not only the model
Changing the model did not solve the main problem.
A new model might write better. Another might reason better. A third might be cheaper for routine tasks. But the model still needed our product knowledge, our company context, our preferences, and the boundaries around what it was allowed to say or do.
A new subscription also meant new memory. The assistant might remember a conversation inside that provider, but it did not give us a durable company context that we controlled. The useful information stayed scattered across accounts and chat histories.
The context window was another limit. Even if we uploaded a large amount of information, the assistant did not automatically know which source was authoritative, which information belonged to which brand, or which personal preferences should stay private.
We needed something above the model layer.
We started with Hermes
We started with Hermes because we wanted an agent operating layer rather than another standalone chat product.
Hermes gave us a way to create profiles, assign tools, add skills, connect external services, schedule work, and keep different agent contexts separate. We could create a LeadCRM agent without making it pretend to be an expert in every other Intricare product.
That distinction became important quickly.
Our products do not share the same positioning, customer questions, product claims, or support processes. A general assistant can easily produce an answer that sounds reasonable but belongs to the wrong product.
With Hermes, we could create specialised profiles such as:
LeadCRM marketing
SalesStack marketing
LinkedFusion support
LeadConnect support
Uberfox support
Arpan email operations
Each profile can have its own instructions, tools, model preferences, knowledge boundaries, and approval rules.
Hermes became the layer that coordinates the work.
Then we added Multica for tasks and execution
Chat is useful, but chat alone is not an operating system.
When an agent finds a customer request, a content opportunity, or an internal follow-up, that work needs an owner, a status, and a place where progress can be tracked.
We added Multica for tasks, projects, approvals, comments, and agent execution. This gave us a practical separation:
Open WebUI: employee conversations
Multica: tasks, projects, approvals, and execution
Hermes: agent intelligence and tools
For example, an email agent can read a customer thread, prepare a reply draft, and create a Multica approval task. The human owner can review and send the message from Gmail. Once the sent message is detected, the related task can be completed automatically.
The agent prepares the work. The person keeps the decision.
That is a much better fit for customer communication, finance-related work, and other operations where an incorrect action can create a real problem.
We added Open WebUI for everyday team access
We wanted employees to have a simple conversational interface. They should not need to understand profiles, gateways, model providers, or agent runtimes to ask a useful question.
Open WebUI became the team-facing chat surface. It gives employees a familiar place to use approved models and agents while Hermes remains the governed execution layer behind it.
This also gave us a cleaner division between ordinary conversation and structured work:
- Open WebUI is where employees chat, ask questions, and build personal working context.
- Multica is where work becomes a task, project, approval, or agent run.
- Hermes is where profiles, skills, tools, and execution rules live.
The interface stays simple. The system behind it can remain sophisticated.
Honcho gave each person private memory
The next problem was personal context.
We wanted the system to remember how a particular employee works without turning that information into company policy. Someone’s preferred writing style, meeting habits, or personal workflow should not become part of the canonical product knowledge base.
We added Honcho as the private memory layer.
Honcho is for things such as:
- employee preferences;
- working style;
- private history;
- recurring personal context; and
- continuity across conversations and tasks where identity is available.
This matters when a team changes models or providers. The user’s useful context should not disappear simply because they stop using one subscription and start using another.
It also matters for privacy. Personal memory and company truth are different categories of information, so they should not be stored in the same place by default.
Hindsight keeps company knowledge grounded
Private memory solved one problem, but it did not solve product accuracy.
For that, we added Hindsight as the approved company and brand knowledge layer.
Hindsight stores the information agents should use when answering questions about a product or company. That can include product documentation, approved claims, policies, source research, legal documents, support guidance, and content rules.
We keep this knowledge separated by brand. LeadCRM has its own knowledge bank. Other products will have their own banks. This prevents an agent from taking a true statement about one product and applying it incorrectly to another.
The distinction is now clear:
Honcho: what this person prefers and remembers
Hindsight: what the company has approved as knowledge
Open WebUI or Multica: what is happening in the current conversation or task
This two-memory design is one of the most important decisions in the system.
A single memory layer is convenient, but it creates difficult questions. Is this fact official? Did an employee mention it once? Is it private? Is it current? Does it apply to every brand?
Separate layers let us answer those questions before the information reaches an agent.
We are not locked into one model provider
The system also gives us freedom at the model layer.
We can connect models in three ways.
Direct provider APIs
A company can connect directly to providers such as OpenAI, Anthropic, Google, or others. This works well when direct provider relationships and model-specific controls are important.
A managed multi-model API
A reseller or aggregation API can provide several model families through one centrally managed connection. This is the approach we currently use. It lets us choose an appropriate model for different tasks without asking every employee to manage several subscriptions, API keys, and billing accounts.
Self-hosted models
For the right workloads, a company can run open-source models in its own environment or adapt a model for a specific domain. This provides more control over the inference environment, although it also introduces infrastructure and evaluation responsibilities.
The important part is that the memory and knowledge layers sit above the model choice. If we change providers, we should not have to rebuild everything the users and the company have already taught the system.
What this has changed for us
The benefit is not that every answer is magically perfect. It is that the system has a better chance of knowing what kind of answer it is supposed to give.
A LeadCRM content agent can use LeadCRM knowledge and content rules. A support profile can focus on customer questions. An email profile can preserve the original recipient context, create drafts, and wait for approval. A task agent can turn a request into structured work instead of leaving it buried in chat.
We still review important outputs. In fact, the system makes review more useful because it can show the context, source material, draft, and next action in one workflow.
We also get more consistent answers across the team. Employees can keep their own working preferences while relying on the same approved company knowledge.
That is what we wanted from AI in the first place: not a tool that sounds confident, but a system that understands which context it should use and where it should stop.
Why we are sharing this
We built this for Intricare because our own work made the gaps obvious. We had multiple subscriptions, fragmented memory, repetitive explanations, and too much manual checking. Building the system forced us to think through identity, knowledge ownership, provider flexibility, permissions, approvals, and the difference between a conversation and a company fact.
We are now also helping companies that face the same problem.
If your team is using AI but still repeating context, checking generic answers, managing too many subscriptions, or worrying about what the system may do with company information, we can help you design and implement an AI operating system around your workforce and workflows.
Contact Intricare Technologies if you want to discuss the architecture, understand how the pieces fit together, or have us implement a similar system for your company.