Mem++ is AIDAChip’s Organizational Memory system: a new kind of AI memory for many authors, shared decisions, and the evidence behind them.
AIDAChip Organizational Memory means preserving the shared context behind an organization’s decisions: who contributed, what was collectively approved, why it mattered, and what was known at the time.
From one conversation to organizational memory
In a conversation, a narrator can update their own facts. Organizational knowledge is spread across independently authored planning notes, issue tickets, meeting minutes, and follow-up messages. People contribute different pieces of evidence, teams decide together, and a later message may propose a change without replacing an earlier approval.
Many authors. Shared decisions. A history that survives. An AI teammate such as Aida can answer a question in one chat; organizational memory requires it to find the relevant decisions across many people, conversations, and original records.
Consider an illustrative network-on-chip (NoC) freeze at 1.2 GHz. The dates, people, and measurements below are a fictional example, not a reported chip program or benchmark case. Worst negative slack (WNS) measures the most severe timing shortfall; an engineering change order (ECO) can improve a path, while multi-corner, multi-mode (MCMM) sign-off still has to cover the full design. These six records come from different authors and sources, not six updates from one person.
Illustrative NoC decision trail — six records, four sources, multiple authors:
[](https://raw.githubusercontent.com/AIDAChip-Inc/mem-plus-plus/main/docs/brand/mem-plus-plus-decision-trail.svg)
The ECO improves one reported timing result, and the netlist regression passes, but the joint review still conditions the Monday freeze on MCMM sign-off. Priya’s later request to restore Friday is a proposal; its newer timestamp does not turn it into an approved change. All six records remain available as evidence.
“What did March 4 timing show?” The illustrative answer is “WNS improved from −120 ps to −35 ps after the clock ECO; MCMM sign-off was still open.” The March 5 positive-slack report and joint approval are outside that time window.
“What NoC freeze did we set, and did the Friday request change it?” The answer is “Monday, pending MCMM sign-off.” Architecture, Physical Design, and Verification agreed in the March 5 review. The March 6 Friday request did not replace that joint decision. This is the paper’s supersession problem: a later record matters only if it actually replaces the earlier decision, not merely because it is newer.
Mem++ preserves and retrieves the source records; the answering model reasons over their contents. The measurements, approval, and condition belong to this fictional example. Mem++ does not measure timing, perform sign-off, create approvals, or enforce an organization’s decision-making rules.
Keep the evidence before interpreting it
Memory systems often transform incoming information into summaries, extracted facts, or updated records. Those representations can be useful, but deciding what to retain at ingestion can also remove context that a later question needs.
Mem++ takes a non-destructive approach. It retains original documents and earlier versions in an append-only store, together with dates and author tags when available. Each new document requires a sentence embedding, but no generative-model call at write time. Empty documents are skipped, as are exact duplicates of active records within the same author scope.
This separates storing evidence from interpreting it. The system does not have to predict every future question before deciding what information deserves to survive. Full document text remains available even when the text used to create an embedding must be truncated.
The result is a memory design centered on inspectable source material. When an answer depends on a changed decision, the earlier record can still be retrieved and presented to the answering model. An optional consolidation pass can link near-duplicate or conflicting versions and mark a record as superseded without deleting it; it was disabled for the headline benchmark results below.
Ask a question, then select the relevant history
When a question arrives, Mem++ uses multi-index fusion: lexical search finds exact terms, semantic search finds paraphrases, and author-tag search contributes where tags are available. Their rankings are combined by weighted reciprocal rank fusion. For an “as of” question, temporal scoping applies the date cutoff before ranking; a recency reserve keeps room for the latest eligible records.
In the NoC example, a March 4 cutoff keeps the Architecture proposal and the Physical Design timing reports eligible, while excluding the later regression result, joint review, and Friday follow-up request. Those original records let the answering model describe what was known before the teams met.
The answer model receives original evidence with dates and interprets the relevant records. This does not make every temporal question automatic or guarantee that conflicting documents will always be resolved correctly. It does preserve the material the model needs to reason about those questions.
The headline Mem++ evaluations use this approach with optional consolidation disabled. The results below refer to the base Mem++ system.
What was known then, what changed, and why
Past context — what was known then? An “as of” cutoff limits the evidence to documents available by the requested date. That helps prevent a later revision from leaking into a historical answer. Bi-temporal questions go further: the date a document was recorded and the date a decision takes effect can differ. Date filtering alone does not solve that reasoning task.
Supersession — what replaced what? A new decision can replace an earlier one without erasing it, but a newer proposal is not automatically a replacement. In the NoC example, the jointly approved Monday freeze and the later Friday request have different meanings in their source documents. Preserving both lets the answering model compare the evidence and explain the history. Mem++ supplies the records; it does not guarantee that the model will identify the authoritative decision in every case.
Decision provenance, justification chains, and audit replay — what supports the answer? Dated original documents preserve who contributed, what was approved, and the conditions behind a change. An answering model can use them to trace a justification chain or reconstruct the evidence available at a past date. This is inspectable evidence for reasoning, rather than a claim of certified audit compliance.
Leading results in the reported evaluations
The Mem++ paper[^1] evaluates organizational and conversational memory. In the reported gpt-4.1-mini comparisons, Mem++ leads the evaluated external memory systems on OrgMemBench overall and achieves the highest average LLM-judge score on LoCoMo.
Organizational Memory
The table below ranks Mem++ and the external memory systems in the paper’s gpt-4.1-mini comparison on OrgMemBench by overall LLM-judge score. Scores are on a 0–100 scale, higher is better. Overall scores include the reported standard deviation across runs. Supersession and audit replay illustrate two specific capabilities; the overall score covers all six question categories.
[](https://raw.githubusercontent.com/AIDAChip-Inc/mem-plus-plus/main/docs/brand/mem-plus-plus-orgmem-results-v3.svg)
Source and scope: Mem++ paper, OrgMemBench evaluation. Reported gpt-4.1-mini LLM-judge scores; comparison here is with external memory-system baselines. The full paper also reports RAG and Memg++ results. These are author-reported research results, not an independently reproduced or live leaderboard.
Conversational Agent Memory
The paper also evaluates conversational memory. Mem++ achieves the highest average LLM-judge score among the evaluated methods on LoCoMo, using gpt-4.1-mini. The following table compares Mem++ with external memory systems reported in the paper:
[](https://raw.githubusercontent.com/AIDAChip-Inc/mem-plus-plus/main/docs/brand/mem-plus-plus-locomo-results-v3.svg)
Source and scope: Mem++ paper, LoCoMo evaluation. Reported average LLM-judge score with gpt-4.1-mini; external baseline results are attributed in the paper to Nan et al. (2025). This is an average-score comparison, not a claim about every question category.
On OrgMemBench, Mem++ leads the strongest external memory-system baseline by 13.1 points with gpt-4.1-mini. This is a score difference within the paper’s comparison, not a percentage improvement.
These results support a practical research direction: preserving documents and selecting evidence at question time can be a strong foundation for agent memory without generative processing on every write.
About the evaluation: OrgMemBench contains 443 dated artifacts from one synthetic organization and 73 questions with 247 graded facets. It tests organizational history in a controlled setting; real-world deployments and author-aware retrieval need further evaluation.
Build on the original record
We’re open-sourcing Mem++ so the community can use it, improve it, and build agents that preserve the evidence behind shared decisions.
For teams developing complex chips, timing reports, ECOs, verification results, and sign-off decisions make preserving context especially relevant. The research asks a broader question: how can agents remain useful as organizational knowledge changes? Mem++ offers one answer—preserve the record, retrieve it in context, and interpret it when needed.
Explore the open-source repository or read the research paper.
---
[^1]: The Mem++ paper is joint research by AIDAChip and The University of Texas at Austin. We are open-sourcing the work for the community.