raminderpalsingh.com
fred project mascot

fredv0.4.1

Uniquely positioned for the secure, private creation and verification of data packages in pre-clinical life sciences. fred reasons over your own data and its own searches of public databases, on a free open-weight model on your own hardware; keeps every run as knowledge you own; and plugs into your workflows and agentic frameworks as a locked vendor service.

fred finds and checks the evidence, and shows where every fact came from. It does not judge the science: your scientists do.

fred building and verifying pre-clinical data packages, privately f fred builds and verifies pre-clinical data packages, privately 1. Data package creation as in Compound Insights Your data README, .md, .json PubMed UniProt PubChem and more public sources Data package claim PMID claim DOI claim UniProt claim yours ? no data found TRACEABLE "Every claim gets a citation. Where the evidence runs out, I say so." 2. Data package verification as in the SIM Framework, and continuously in MAVIS Data package Sufficient! Not sufficient! same evidence, both sides, five times over citations curated retrieved neither your committee decides "Both sides, every citation checked. I don't decide; you do." All on your own machine, on a free open-weight model. Your data and fred's reasoning never go to a model provider. Not air-gapped: fred's searches of public databases do leave the machine. public databases searches only
Figure 1. The two jobs fred does. Left: building a data package from your own data and public sources, with every claim traced and gaps stated, as Compound Insights does. Right: verifying a package by arguing each claim both ways and checking every citation, with the decision left to people, as the SIM Framework does at a decision point and MAVIS does continuously. Both run on your own machine; only fred's database searches leave it.
open weightsA free model from Hugging Face. No cloud model and no per-call fees.
32 GB VRAMThe model and the service run together on one machine.
sourced claimsIdentifiers are traced to a database reply, or flagged.
every run keptResults, sources and decisions go into a data layer your applications can read.
locked servicePlugs into your workflows and agentic frameworks over plain HTTP. Nothing inside changes.

What fred is

A research engine that looks things up, shows its sources, and remembers what it found.

Pre-clinical decisions turn on specific values: an accession, a compound identifier, a PubMed ID, a molecular weight, a sequence length. General-purpose chat models write fluent answers whose identifiers can be wrong in ways that look right. fred is built the other way round: the answer is assembled from what its own database and literature searches actually return, and anything it cannot trace is marked.

fred reasons; it does not judge the science. It establishes what the evidence says, where each value came from, where sources disagree, and what it could not verify. Whether a target is worth pursuing, whether a compound is safe enough, or whether a package is ready to advance are judgments for scientists, and every application built on fred leaves them with people.

fred is a ReAct reasoning engine delivered as a containerized HTTP service. An application gives fred two things. The first is its context: its own data as Markdown and JSON files, such as internal reports, assay summaries, curated records and project notes, with a README that lists each file, says where it came from and what not to rely on. The context is the critical background for the analysis: it tells fred what the organization already knows and what the question is really about. The second is the question, which names that context. fred reads the context first, decides what to check, queries public databases one step at a time, and returns an answer in which each identifier is traced to the context or to a database reply, together with a record of every lookup. It is a vendor service: one installation serves many applications, such as Compound Insights and SIM, which call it and never change it, and it is built around none of them.

Nothing in fred fails silently. A lookup that fails says so and why. A model call that runs out of room is retried or fails the run; it is never passed on as a complete answer. An identifier fred could not verify is flagged in the answer itself, where the reader will see it, not only in a side report. Even a failure to record a run is counted. A scientist relying on fred should never have to guess whether something went wrong.

Version 0.4 is designed around four commitments. Version 0.4.1 makes your own data a first-class input: a context directory of Markdown and JSON files, indexed by a README, that fred checks, seals, reads as labeled background in every question that names it, grounds its answers against, and keeps for your later questions.

Your application 1. uploads its context 2. asks a question that names that context 6. reads the answer, where each fact came from, and what fred could not verify Your context critical background README.md: the index, what each file is, its source and its caveats .md and .json files of your own data checked, then sealed fred v0.4.1: a ReAct loop each step decided by the local model 3. THOUGHT what does the context say? what to check? 4. ACTION call one of ten tools 5. OBSERVATION read the reply, record it Grounding check on the answer each identifier traced to a lookup or to your context, or flagged; conflicts between them reported Public databases UniProt (proteins) PubChem (compounds) PubMed (literature) Open Targets Reactome (pathways) Web search over the internet Local model open weights, 8-bit, free from Hugging Face starts on demand, stops when idle on your machine; no cloud Data layers: everything kept what fred generated: answers, tool calls, sources, loops, retries, corrections what you sent: context files and README each application reads only its own records Release checks same questions, same day, against the previous release Reuse grounded history only, with its sources, within one application; default set by measurement context question answer query reply steps kept
Figure 2. fred v0.4.1. Your application uploads its context, its own Markdown and JSON files with a README that indexes them, and fred checks and seals it. Each question names that context, which the model reads first as labeled background before it searches. The loop of thought, action and observation runs on a free open-weight model on your own machine; the lookups go to public databases. Before an answer leaves, a grounding check traces its identifiers to a lookup or to your context and reports where the two disagree. What fred generates and what you sent are both kept, and each application reads only its own. Reuse of that history (dashed) stays within one application. Every release is tested against the one before it.

What fred is for

Building and verifying pre-clinical data packages, privately.

Pre-clinical decisions are made on data packages: a compound dossier, a target assessment, the nomination package a committee reads before committing a molecule to regulated toxicology. Each package has to be assembled from many sources with every claim traceable, and each has to be challenged before anyone commits money, time and animals to it. fred is uniquely positioned to do both jobs on an organization's own hardware. Its first applications show it doing each, and a third, MAVIS, keeps the verification going after a package is written, because evidence changes.

Data package creation: Compound InsightsCompound Insights builds cited compound dossiers. fred synthesizes narrative claims from the primary literature, such as binding mode, kinase class, pharmacokinetics, regulatory history and off-target mechanisms, and the engine joins them with structured records from ChEMBL, UniProt, ClinicalTrials.gov, Reactome, Monarch and Open Targets. Every claim fred contributes carries a PubMed ID or DOI. For an investigational molecule with only three PubMed records, much of the dossier is an explicit "no data found" rather than an inference: the test is whether the engine declines where the evidence runs out. Data package verification: SIM FrameworkSIM takes the package for a Development Candidate decision and argues each of its claims both ways: the strongest case that the evidence is sufficient, the strongest case that it is not, and an adjudication that checks every citation and may decline to rule, repeated five times with disagreement reported rather than averaged away. fred produces every argument and every adjudication. SIM tags each citation as curated, retrieved or neither, keeps every failed repetition with its reason, and makes no go/no-go call: a human committee decides. Continuous verification: MAVISMAVIS, an evidence guard in development, keeps a standing watch on the claims AI tools put into a program's research files. It sends fred the program's own Markdown and JSON files with their README and has fred check every claim that names an identifier or a value with a unit. It checks again when a file changes, on a schedule, or when a source may have changed, and tells the claim's owner, with the evidence, when a claim cannot be traced, conflicts with its source, or stops holding. People decide what to do; MAVIS records the decision.

Why fred, for this work

Private is not air-gapped. fred's database and literature searches leave the machine while it reasons, as SIM and MAVIS state on their own pages.

Built for pre-clinical science decisions

An answer a scientist can check, value by value.

A plausible sentence carrying a wrong identifier is worse than no answer, because it looks finished. A language model's own recall of accessions, compound identifiers and PubMed IDs is unreliable, so fred does not depend on it. It looks values up, checks the answer against what it found, and tells you where the two disagree.

CapabilitySpecification
evidence toolsTen tools: UniProt protein lookup, PubChem compound search and properties, PubMed publication search, Open Targets, Reactome pathways, a Lipinski rule-of-five calculator, web search, and two workspace file tools.
your context (0.4.1)An application uploads its own data into a context directory: Markdown files of free text and JSON files in a structured record envelope, within limits on file size and count. A required README.md, in a fixed template, lists and categorizes every file: format, category, source type, date, fingerprint, description, source and caveats. fred checks the directory as a whole when it is sealed and refuses a malformed one with a named error, before any model call. Every question that names the directory gets it as labeled background, README first.
grounding checkEvery completed answer is checked: each identifier is compared with the tool replies and the context. Each claim is labeled by where it came from: a tool reply, the question, the context, or unsourced. Where the context and a database disagree, fred reports the conflict rather than choosing between them. A mismatch gets one correction turn that lists the near-matches fred did see. Anything still unsupported is flagged in the answer with an "Unverified:" note and in a machine-readable verdict.
provenance recordFor every tool call: the tool, its arguments, start time, duration, outcome (found, found nothing, or failed, with the reason), and the size and SHA-256 fingerprint of the reply.
structured answersA question can require a JSON answer of a given shape, so an application can parse it directly; conformance is measured in every release check.
honest failuresEvery lookup reports found, found nothing, or failed with the reason. Unverified identifiers are flagged in the answer text. A truncated or runaway model call is retried or fails the run, never returned as complete.
honest gapsWhen the evidence, searched or in the context, does not establish a value, the expected answer says so rather than filling the gap. Release checks include questions built to test exactly this.

fred is a research aid. Treat its answers as evidence with sources, not as final answers: the record is there so a scientist can see how an answer was reached and check it.

Open-weight models on hardware you can buy

The model, the question and the reasoning stay on your machine.

Pre-clinical questions often carry unpublished hypotheses, and the data sent with them is often proprietary. fred's reasoning runs on a free open-weight model on your own hardware, so those questions, your data and the model's reasoning about them never go to a model provider. Search terms fred derives from them do reach the public databases it queries. There are no per-call fees, and the model cannot change underneath a result: its weights are pinned to a published revision.

CapabilitySpecification
reference modelQwen3.8 27B, a 27-billion-parameter open-weight model released under Apache 2.0 and free to download from Hugging Face (Qwen/Qwen3.8-27B), run at 8-bit precision. 8-bit is chosen over smaller 4-bit builds to keep as much of the model's reasoning as the hardware allows, accepting slower generation: for a scientific decision, answer quality comes before speed.
model choicefred is model-agnostic. The model and its server are configuration, not code, so another open-weight model can be used as better ones are released. Each release is verified with the reference model.
no hidden layerfred builds every request itself: its instructions, tool definitions and messages reach the model exactly as sent. In testing, a model runtime that silently rewrote fred's tool definitions changed how the model behaved, so fred does not depend on any runtime's undisclosed processing.
hardware32 GB of VRAM for the model at 8-bit. The model, its server and fred run on the same machine.
model serverAny local server with an OpenAI-compatible API. The reference setup starts the model on the first question and unloads it after five minutes idle, so the machine is free between uses. No dependency on Ollama or any other single runtime.
local by constructionfred has no built-in model or server address and will not start until both are named. It never calls a cloud model.
loop controlEvery model call is streamed. A repeating loop is detected, the generation is stopped, and the turn is retried with a direct instruction, at most twice per question. Repetition settings are adopted only when measured to help.
deadlinesEach model call has a total deadline (600 s by default). The model's reasoning is kept out of the answer text.
speedOne question at a time; a question that takes several steps takes minutes. Faster decoding, such as speculative decoding with a small draft model, is adopted only where it is measured to pay without costing accuracy.
portabilityA Docker container, with its data layer in a Docker volume, for macOS and Linux hosts.

Tacit knowledge, kept and reused

What fred observed, prioritized and decided, run after run.

An experienced researcher knows more than the answers they have given. They know which sources were reliable, which searches came back empty, which identifiers are easy to confuse, and which questions needed a second look. That is tacit knowledge, and it is usually lost when the person moves on.

fred keeps the equivalent from every question it answers: which sources answered, which searches came back empty, which identifiers had to be corrected, and which source each claim rests on. The store grows with every run and becomes more valuable as it does: a readable, auditable history of the evidence behind every past answer, held on your own hardware and belonging to you rather than to a model provider. Each application reads its own, and grounded history is offered back to new questions, switched on wherever it is shown to improve the answers.

What is kept

RecordContents
runsThe question, the context it named, the answer, and how the run ended.
your context (0.4.1)Every context directory, its README and files kept whole, indexed by the README's fields, and linked to the runs that used it. It is a second data layer beside what fred generates.
tool callsEvery call's arguments and the reply as the model saw it, with its fingerprint.
eventsLoops found, retries, correction turns and limits reached.
claimsEach claim in an answer, linked to where it first appeared and labeled: from a tool reply, from the question's own text, from the context, or unsourced.
reasoningThe model's reasoning, in its own table, labeled as unverified model text and never eligible for reuse.

Access and reuse

CapabilitySpecification
read accessRead-only routes return an application's own runs, using its tenant token. Another application's runs answer as not found.
retentionRecords are kept with no automatic expiry. Only the operator exports or deletes them; applications cannot delete.
test isolationTest runs are stored apart, returned to no application, and never used as history.
capture safetyThe record is written after the answer is fixed and never changes it; a failed write is counted, not allowed to affect the run. No credential value is stored.
reuseOnly grounded items, whose facts came from a tool reply or from that application's context, are offered to a new question, each with its original run and sources. A switch per question controls it, and its default is set by measured results on related questions.

Reuse can also repeat a past mistake with more confidence, which is why only grounded items with their sources qualify, and why its default follows the measurements rather than the expectation.

Built to drop into your workflows

A locked vendor service that agent frameworks, harnesses and pipelines call, and never change.

Research organizations increasingly run their work through orchestrated workflows: agent frameworks, evaluation harnesses and pipelines that retrieve, reason, check and compile across many steps. fred is built to be one dependable step inside them. It ships as a sealed, versioned container with a small and stable HTTP contract, so an orchestrator treats it the way it treats any vendor service: it calls fred, reads a structured result, and never needs to know or change what is inside.

CapabilitySpecification
locked serviceEach release is an immutable container image with a version and a digest. An application changes one thing to use fred: the address it calls. Upgrading fred is a deliberate, recorded event, never a side effect of someone else's work.
stable contractA wire version on every request and response, and changes within a version are additive only. The tool registry is published at its own route, so a harness can pin it and notice any change before it trusts a result.
bring your contextA workflow uploads the evidence it already holds as a sealed context directory, once, and names it in each question (from 0.4.1), so fred reasons over the organization's own records as well as public sources.
any frameworkPlain HTTP and JSON: submit a question, poll, read the result. No SDK, no client library and no language lock-in. Any agent framework or workflow engine that can make an HTTP call can use fred as a tool or as a reasoning step.
results a harness can act onThe answer comes with a machine-readable grounding verdict, a provenance record for every tool call with fingerprints, a fixed outcome vocabulary, and an optional required JSON shape. A harness can validate an answer, tag each citation by where it came from, and retry or label a weak result without parsing prose.
control per runA step budget (1 to 200 turns, 30 by default), the in-answer "Unverified:" note on or off, and history reuse on or off. A busy service answers at once with the run in progress, so a harness can schedule rather than wait blind.
identity and auditEach application or agent gets its own tenant token. Its runs are kept and can be read back afterward, so every step a workflow delegated to fred can be audited.
data stays infred listens only on the machine's own loopback address, beside the workflow that calls it. Only the database lookups leave the machine.

Compound Insights, SIM and MAVIS, described under What fred is for, are independent applications that share no code with each other and use fred exactly this way. MAVIS shows the harness pattern most clearly: it runs fred in a standing loop over a program's files, compares each result with the last, counts a claim as holding only when a public source confirmed it, and raises an alert only when a second check agrees, because fred's answers can vary from run to run.

How it is verified

Paired runs, a measured noise floor, and claims that name their cases.

fred is not deterministic, so one run proves nothing and a small difference may be noise. Every release is measured against the one before it with paired runs: the same questions on both, back to back on the same day, in alternating order, with the same independent judge and the same verification script. A difference counts only if it is larger than the difference measured between two identical services. Thresholds come from measured spread, never from a guess.

The questions are fixed before fred sees them, with expected answers taken from database replies saved with their fingerprints. They include registered retrieval questions (a protein in UniProt, an article in PubMed, a compound in PubChem) and questions of the kind applications send: sets of related PubMed abstracts with a required JSON answer, some needing database lookups as well and some asking for a value the text does not establish.

CheckPass condition
toolsEvery tool reachable, and each kind of tool failure reported as a failure, not a quiet pass.
citationsNo identifier the databases never returned, none that fails to resolve, and every required citation present.
judgeNo answer scored wrong by the independent judge, which is calibrated first on deliberately corrupted answers.
baselineResults match the stored baseline cell by cell; the comparison refuses if the questions or the yardstick have changed.
answer shapeAnswers follow the required JSON shape, and identifiers outside the question's own text and the tool replies are counted.
contextEvery malformed context directory is refused with a named error before any model call, and cases with instructions planted in context files test whether a run can be steered.
false alarmsThe grounding check never flags an item that is in the question's own text.
CheckPass condition
local onlyEvery model call fred makes is served by the local model; any other call fails the run.
loopsLoops, runaway reasoning and failed runs are counted per question and compared with the previous release.
timeTime per question, per step and per model call, inside every budget.
memoryThe machine's memory is watched throughout; pressure or swap growth outside a model load stops the run.
diskThe model server's cache stays bounded and is emptied at each start.
CheckPass condition
complete recordsEvery run leaves a complete record, readable by its own application and by no other.
no credentialsThe store is scanned for credential values; none may be found.
no side effectAnswers and times with capture on stay within the noise floor of the previous release.
sizeThe store's size per run and in total is recorded after every batch.
reuseOn related questions, accuracy with history is compared with accuracy without it; reused items that were not grounded must number zero; dependence on the order of earlier runs is reported.

The judge belongs to the test suite, not to fred, and runs on a cloud model; fred itself never calls one. Each published result names the questions it rests on.

Requirements

One machine, a container runtime, and the model.

ItemRequirement
GPU memory32 GB of VRAM for the reference model at 8-bit.
modelQwen3.8 27B at 8-bit, free from Hugging Face, served by a local OpenAI-compatible model server.
runtimeDocker, on macOS or Linux.
diskAbout 30 GB for the model weights, plus room for the data layer, which grows with use.
networkOutbound HTTPS to the public databases and a web search provider.
keysAn NCBI API key for PubMed and a web search API key, supplied through the environment. No model API key.

API

One service, many applications. Plain HTTP and JSON on a loopback address, no SDK.

RouteWhat it does
GET /v1/healthLiveness, and the run in progress, or null when idle.
GET /v1/toolsThe tool registry: name, description and argument schema for each tool.
POST /v1/runsSubmit a question. Answers 202 with a run identifier, or 503 with the run in progress.
GET /v1/runs/{run_id}Status, then the answer, the grounding verdict and the record of every tool call.
GET /v1/records/runsYour application's kept runs, with your tenant token in the X-Fred-Tenant-Token header.
GET /v1/records/runs/{run_id}One kept run in full: tool calls and replies, events and labeled claims.
GET /v1/records/runs/{run_id}/reasoningThat run's model reasoning, labeled as unverified.
GET /v1/records/statusCounts only.

Send "wire_version": "1.2" on every request and check that the response carries the same value. Changes within a wire version are additive: no existing field changes meaning. The "Unverified:" note is in the answer by default, and a request field turns it off for applications that read the verdict themselves. From 0.4.1, at wire version 1.3, further routes create a context directory, upload files into it, seal it, read it back, retire it and delete it, each checked against the tenant token, and a run names a sealed directory by its id. Check that the service reports wire 1.3 before sending context.

Known limits

In plain words.

fred reasons; it does not judge the science

fred checks whether claims trace to evidence and still agree with it. It does not judge whether the science is right or whether a molecule should advance. Applications built on it, such as SIM and MAVIS, put the evidence in front of the people who decide; nothing fred returns is a recommendation to act.

fred is not deterministic

Ask the same question twice and you may get different lookups and different wording.

Measured on a focused set of questions

Release checks cover registered retrieval questions and abstract-based questions of the applications' shape. A result on them is evidence about questions of that shape, not proof on every question.

The model is local; the lookups are not

The model and its reasoning stay on your machine. The search terms fred sends to the public databases and the web search provider leave it.

Minutes, not seconds

A question that takes several steps takes minutes, and the first question after an idle period also waits for the model to load. fred answers one question at a time.

Every run is kept

The data layers hold each application's questions, context files and answers until the operator deletes them.

Separation rests on tokens

The service listens only on the machine's own loopback address, without authentication. Applications are kept apart in the data layer by tenant tokens: a process that obtains an application's token can read that application's records.

Your context can mislead

fred treats your files as reference data, not instructions, and tells the model so. That reduces, and does not remove, the chance that text in a file steers a run. A wrong value in your context is grounded under the context's own label, and fred cannot detect secrets you include.

Reuse can repeat a mistake

An earlier answer can be wrong in a way fred's own checks accepted. That is why reuse admits only grounded items with their sources, and why its default is set by measurement.