Uniquely positioned for the secure, private creation and verification of data packages in pre-clinical life sciences. fred reasons over your own data and its own searches of public databases, on a free open-weight model on your own hardware; keeps every run as knowledge you own; and plugs into your workflows and agentic frameworks as a locked vendor service.
fred finds and checks the evidence, and shows where every fact came from. It does not judge the science: your scientists do.
A research engine that looks things up, shows its sources, and remembers what it found.
Pre-clinical decisions turn on specific values: an accession, a compound identifier, a PubMed ID, a molecular weight, a sequence length. General-purpose chat models write fluent answers whose identifiers can be wrong in ways that look right. fred is built the other way round: the answer is assembled from what its own database and literature searches actually return, and anything it cannot trace is marked.
fred reasons; it does not judge the science. It establishes what the evidence says, where each value came from, where sources disagree, and what it could not verify. Whether a target is worth pursuing, whether a compound is safe enough, or whether a package is ready to advance are judgments for scientists, and every application built on fred leaves them with people.
fred is a ReAct reasoning engine delivered as a containerized HTTP service. An application gives fred two things. The first is its context: its own data as Markdown and JSON files, such as internal reports, assay summaries, curated records and project notes, with a README that lists each file, says where it came from and what not to rely on. The context is the critical background for the analysis: it tells fred what the organization already knows and what the question is really about. The second is the question, which names that context. fred reads the context first, decides what to check, queries public databases one step at a time, and returns an answer in which each identifier is traced to the context or to a database reply, together with a record of every lookup. It is a vendor service: one installation serves many applications, such as Compound Insights and SIM, which call it and never change it, and it is built around none of them.
Nothing in fred fails silently. A lookup that fails says so and why. A model call that runs out of room is retried or fails the run; it is never passed on as a complete answer. An identifier fred could not verify is flagged in the answer itself, where the reader will see it, not only in a side report. Even a failure to record a run is counted. A scientist relying on fred should never have to guess whether something went wrong.
Version 0.4 is designed around four commitments. Version 0.4.1 makes your own data a first-class input: a context directory of Markdown and JSON files, indexed by a README, that fred checks, seals, reads as labeled background in every question that names it, grounds its answers against, and keeps for your later questions.
Building and verifying pre-clinical data packages, privately.
Pre-clinical decisions are made on data packages: a compound dossier, a target assessment, the nomination package a committee reads before committing a molecule to regulated toxicology. Each package has to be assembled from many sources with every claim traceable, and each has to be challenged before anyone commits money, time and animals to it. fred is uniquely positioned to do both jobs on an organization's own hardware. Its first applications show it doing each, and a third, MAVIS, keeps the verification going after a package is written, because evidence changes.
Private is not air-gapped. fred's database and literature searches leave the machine while it reasons, as SIM and MAVIS state on their own pages.
An answer a scientist can check, value by value.
A plausible sentence carrying a wrong identifier is worse than no answer, because it looks finished. A language model's own recall of accessions, compound identifiers and PubMed IDs is unreliable, so fred does not depend on it. It looks values up, checks the answer against what it found, and tells you where the two disagree.
| Capability | Specification |
|---|---|
| evidence tools | Ten tools: UniProt protein lookup, PubChem compound search and properties, PubMed publication search, Open Targets, Reactome pathways, a Lipinski rule-of-five calculator, web search, and two workspace file tools. |
| your context (0.4.1) | An application uploads its own data into a context directory: Markdown files of free text and JSON files in a structured record envelope, within limits on file size and count. A required README.md, in a fixed template, lists and categorizes every file: format, category, source type, date, fingerprint, description, source and caveats. fred checks the directory as a whole when it is sealed and refuses a malformed one with a named error, before any model call. Every question that names the directory gets it as labeled background, README first. |
| grounding check | Every completed answer is checked: each identifier is compared with the tool replies and the context. Each claim is labeled by where it came from: a tool reply, the question, the context, or unsourced. Where the context and a database disagree, fred reports the conflict rather than choosing between them. A mismatch gets one correction turn that lists the near-matches fred did see. Anything still unsupported is flagged in the answer with an "Unverified:" note and in a machine-readable verdict. |
| provenance record | For every tool call: the tool, its arguments, start time, duration, outcome (found, found nothing, or failed, with the reason), and the size and SHA-256 fingerprint of the reply. |
| structured answers | A question can require a JSON answer of a given shape, so an application can parse it directly; conformance is measured in every release check. |
| honest failures | Every lookup reports found, found nothing, or failed with the reason. Unverified identifiers are flagged in the answer text. A truncated or runaway model call is retried or fails the run, never returned as complete. |
| honest gaps | When the evidence, searched or in the context, does not establish a value, the expected answer says so rather than filling the gap. Release checks include questions built to test exactly this. |
fred is a research aid. Treat its answers as evidence with sources, not as final answers: the record is there so a scientist can see how an answer was reached and check it.
The model, the question and the reasoning stay on your machine.
Pre-clinical questions often carry unpublished hypotheses, and the data sent with them is often proprietary. fred's reasoning runs on a free open-weight model on your own hardware, so those questions, your data and the model's reasoning about them never go to a model provider. Search terms fred derives from them do reach the public databases it queries. There are no per-call fees, and the model cannot change underneath a result: its weights are pinned to a published revision.
| Capability | Specification |
|---|---|
| reference model | Qwen3.8 27B, a 27-billion-parameter open-weight model released under Apache 2.0 and free to download from Hugging Face (Qwen/Qwen3.8-27B), run at 8-bit precision. 8-bit is chosen over smaller 4-bit builds to keep as much of the model's reasoning as the hardware allows, accepting slower generation: for a scientific decision, answer quality comes before speed. |
| model choice | fred is model-agnostic. The model and its server are configuration, not code, so another open-weight model can be used as better ones are released. Each release is verified with the reference model. |
| no hidden layer | fred builds every request itself: its instructions, tool definitions and messages reach the model exactly as sent. In testing, a model runtime that silently rewrote fred's tool definitions changed how the model behaved, so fred does not depend on any runtime's undisclosed processing. |
| hardware | 32 GB of VRAM for the model at 8-bit. The model, its server and fred run on the same machine. |
| model server | Any local server with an OpenAI-compatible API. The reference setup starts the model on the first question and unloads it after five minutes idle, so the machine is free between uses. No dependency on Ollama or any other single runtime. |
| local by construction | fred has no built-in model or server address and will not start until both are named. It never calls a cloud model. |
| loop control | Every model call is streamed. A repeating loop is detected, the generation is stopped, and the turn is retried with a direct instruction, at most twice per question. Repetition settings are adopted only when measured to help. |
| deadlines | Each model call has a total deadline (600 s by default). The model's reasoning is kept out of the answer text. |
| speed | One question at a time; a question that takes several steps takes minutes. Faster decoding, such as speculative decoding with a small draft model, is adopted only where it is measured to pay without costing accuracy. |
| portability | A Docker container, with its data layer in a Docker volume, for macOS and Linux hosts. |
What fred observed, prioritized and decided, run after run.
An experienced researcher knows more than the answers they have given. They know which sources were reliable, which searches came back empty, which identifiers are easy to confuse, and which questions needed a second look. That is tacit knowledge, and it is usually lost when the person moves on.
fred keeps the equivalent from every question it answers: which sources answered, which searches came back empty, which identifiers had to be corrected, and which source each claim rests on. The store grows with every run and becomes more valuable as it does: a readable, auditable history of the evidence behind every past answer, held on your own hardware and belonging to you rather than to a model provider. Each application reads its own, and grounded history is offered back to new questions, switched on wherever it is shown to improve the answers.
| Record | Contents |
|---|---|
| runs | The question, the context it named, the answer, and how the run ended. |
| your context (0.4.1) | Every context directory, its README and files kept whole, indexed by the README's fields, and linked to the runs that used it. It is a second data layer beside what fred generates. |
| tool calls | Every call's arguments and the reply as the model saw it, with its fingerprint. |
| events | Loops found, retries, correction turns and limits reached. |
| claims | Each claim in an answer, linked to where it first appeared and labeled: from a tool reply, from the question's own text, from the context, or unsourced. |
| reasoning | The model's reasoning, in its own table, labeled as unverified model text and never eligible for reuse. |
| Capability | Specification |
|---|---|
| read access | Read-only routes return an application's own runs, using its tenant token. Another application's runs answer as not found. |
| retention | Records are kept with no automatic expiry. Only the operator exports or deletes them; applications cannot delete. |
| test isolation | Test runs are stored apart, returned to no application, and never used as history. |
| capture safety | The record is written after the answer is fixed and never changes it; a failed write is counted, not allowed to affect the run. No credential value is stored. |
| reuse | Only grounded items, whose facts came from a tool reply or from that application's context, are offered to a new question, each with its original run and sources. A switch per question controls it, and its default is set by measured results on related questions. |
Reuse can also repeat a past mistake with more confidence, which is why only grounded items with their sources qualify, and why its default follows the measurements rather than the expectation.
A locked vendor service that agent frameworks, harnesses and pipelines call, and never change.
Research organizations increasingly run their work through orchestrated workflows: agent frameworks, evaluation harnesses and pipelines that retrieve, reason, check and compile across many steps. fred is built to be one dependable step inside them. It ships as a sealed, versioned container with a small and stable HTTP contract, so an orchestrator treats it the way it treats any vendor service: it calls fred, reads a structured result, and never needs to know or change what is inside.
| Capability | Specification |
|---|---|
| locked service | Each release is an immutable container image with a version and a digest. An application changes one thing to use fred: the address it calls. Upgrading fred is a deliberate, recorded event, never a side effect of someone else's work. |
| stable contract | A wire version on every request and response, and changes within a version are additive only. The tool registry is published at its own route, so a harness can pin it and notice any change before it trusts a result. |
| bring your context | A workflow uploads the evidence it already holds as a sealed context directory, once, and names it in each question (from 0.4.1), so fred reasons over the organization's own records as well as public sources. |
| any framework | Plain HTTP and JSON: submit a question, poll, read the result. No SDK, no client library and no language lock-in. Any agent framework or workflow engine that can make an HTTP call can use fred as a tool or as a reasoning step. |
| results a harness can act on | The answer comes with a machine-readable grounding verdict, a provenance record for every tool call with fingerprints, a fixed outcome vocabulary, and an optional required JSON shape. A harness can validate an answer, tag each citation by where it came from, and retry or label a weak result without parsing prose. |
| control per run | A step budget (1 to 200 turns, 30 by default), the in-answer "Unverified:" note on or off, and history reuse on or off. A busy service answers at once with the run in progress, so a harness can schedule rather than wait blind. |
| identity and audit | Each application or agent gets its own tenant token. Its runs are kept and can be read back afterward, so every step a workflow delegated to fred can be audited. |
| data stays in | fred listens only on the machine's own loopback address, beside the workflow that calls it. Only the database lookups leave the machine. |
Compound Insights, SIM and MAVIS, described under What fred is for, are independent applications that share no code with each other and use fred exactly this way. MAVIS shows the harness pattern most clearly: it runs fred in a standing loop over a program's files, compares each result with the last, counts a claim as holding only when a public source confirmed it, and raises an alert only when a second check agrees, because fred's answers can vary from run to run.
Paired runs, a measured noise floor, and claims that name their cases.
fred is not deterministic, so one run proves nothing and a small difference may be noise. Every release is measured against the one before it with paired runs: the same questions on both, back to back on the same day, in alternating order, with the same independent judge and the same verification script. A difference counts only if it is larger than the difference measured between two identical services. Thresholds come from measured spread, never from a guess.
The questions are fixed before fred sees them, with expected answers taken from database replies saved with their fingerprints. They include registered retrieval questions (a protein in UniProt, an article in PubMed, a compound in PubChem) and questions of the kind applications send: sets of related PubMed abstracts with a required JSON answer, some needing database lookups as well and some asking for a value the text does not establish.
| Check | Pass condition |
|---|---|
| tools | Every tool reachable, and each kind of tool failure reported as a failure, not a quiet pass. |
| citations | No identifier the databases never returned, none that fails to resolve, and every required citation present. |
| judge | No answer scored wrong by the independent judge, which is calibrated first on deliberately corrupted answers. |
| baseline | Results match the stored baseline cell by cell; the comparison refuses if the questions or the yardstick have changed. |
| answer shape | Answers follow the required JSON shape, and identifiers outside the question's own text and the tool replies are counted. |
| context | Every malformed context directory is refused with a named error before any model call, and cases with instructions planted in context files test whether a run can be steered. |
| false alarms | The grounding check never flags an item that is in the question's own text. |
| Check | Pass condition |
|---|---|
| local only | Every model call fred makes is served by the local model; any other call fails the run. |
| loops | Loops, runaway reasoning and failed runs are counted per question and compared with the previous release. |
| time | Time per question, per step and per model call, inside every budget. |
| memory | The machine's memory is watched throughout; pressure or swap growth outside a model load stops the run. |
| disk | The model server's cache stays bounded and is emptied at each start. |
| Check | Pass condition |
|---|---|
| complete records | Every run leaves a complete record, readable by its own application and by no other. |
| no credentials | The store is scanned for credential values; none may be found. |
| no side effect | Answers and times with capture on stay within the noise floor of the previous release. |
| size | The store's size per run and in total is recorded after every batch. |
| reuse | On related questions, accuracy with history is compared with accuracy without it; reused items that were not grounded must number zero; dependence on the order of earlier runs is reported. |
The judge belongs to the test suite, not to fred, and runs on a cloud model; fred itself never calls one. Each published result names the questions it rests on.
One machine, a container runtime, and the model.
| Item | Requirement |
|---|---|
| GPU memory | 32 GB of VRAM for the reference model at 8-bit. |
| model | Qwen3.8 27B at 8-bit, free from Hugging Face, served by a local OpenAI-compatible model server. |
| runtime | Docker, on macOS or Linux. |
| disk | About 30 GB for the model weights, plus room for the data layer, which grows with use. |
| network | Outbound HTTPS to the public databases and a web search provider. |
| keys | An NCBI API key for PubMed and a web search API key, supplied through the environment. No model API key. |
One service, many applications. Plain HTTP and JSON on a loopback address, no SDK.
| Route | What it does |
|---|---|
| GET /v1/health | Liveness, and the run in progress, or null when idle. |
| GET /v1/tools | The tool registry: name, description and argument schema for each tool. |
| POST /v1/runs | Submit a question. Answers 202 with a run identifier, or 503 with the run in progress. |
| GET /v1/runs/{run_id} | Status, then the answer, the grounding verdict and the record of every tool call. |
| GET /v1/records/runs | Your application's kept runs, with your tenant token in the X-Fred-Tenant-Token header. |
| GET /v1/records/runs/{run_id} | One kept run in full: tool calls and replies, events and labeled claims. |
| GET /v1/records/runs/{run_id}/reasoning | That run's model reasoning, labeled as unverified. |
| GET /v1/records/status | Counts only. |
Send "wire_version": "1.2" on every request and check that the response carries the same value. Changes within a wire version are additive: no existing field changes meaning. The "Unverified:" note is in the answer by default, and a request field turns it off for applications that read the verdict themselves. From 0.4.1, at wire version 1.3, further routes create a context directory, upload files into it, seal it, read it back, retire it and delete it, each checked against the tenant token, and a run names a sealed directory by its id. Check that the service reports wire 1.3 before sending context.
In plain words.
fred checks whether claims trace to evidence and still agree with it. It does not judge whether the science is right or whether a molecule should advance. Applications built on it, such as SIM and MAVIS, put the evidence in front of the people who decide; nothing fred returns is a recommendation to act.
Ask the same question twice and you may get different lookups and different wording.
Release checks cover registered retrieval questions and abstract-based questions of the applications' shape. A result on them is evidence about questions of that shape, not proof on every question.
The model and its reasoning stay on your machine. The search terms fred sends to the public databases and the web search provider leave it.
A question that takes several steps takes minutes, and the first question after an idle period also waits for the model to load. fred answers one question at a time.
The data layers hold each application's questions, context files and answers until the operator deletes them.
The service listens only on the machine's own loopback address, without authentication. Applications are kept apart in the data layer by tenant tokens: a process that obtains an application's token can read that application's records.
fred treats your files as reference data, not instructions, and tells the model so. That reduces, and does not remove, the chance that text in a file steers a run. A wrong value in your context is grounded under the context's own label, and fred cannot detect secrets you include.
An earlier answer can be wrong in a way fred's own checks accepted. That is why reuse admits only grounded items with their sources, and why its default is set by measurement.