Aristotle / AI co-scientist for scientific research

An AI co-scientist built for researchers trained to question everything.

Aristotle is Autopoiesis Sciences’ AI co-scientist. I’ve designed it from 0 to 1, turning a black box that picked models on its own into an interface researchers can inspect, question and steer.

Role
Product Designer, 0 → 1
Timeline
Oct 2025 to Present
Team
2 designers + 3 engineers
Company
Autopoiesis Sciences

01 · The Beta

Designing the foundation

Oct 2025 → Feb 2026

Why

Scientists don’t trust AI. And they shouldn’t have to.

Most AI tools are designed to make answers feel effortless. They return a polished response with little indication of how the answer was formed, how reliable the evidence is, or where the researcher should start questioning it.

That model works for everyday questions. Scientific research demands something different.

Researchers need to be able to:

  • See what the AI is doing.Understand how it approached the question and what evidence shaped the response.
  • Question the answer.Identify uncertainty, weak claims, and gaps instead of accepting a confident response at face value.
  • Work with rigor.Move from an initial hypothesis to evidence, verification, and deeper investigation without leaving the research workflow.

The challenge wasn’t simply making Aristotle more accurate.

It was designing an interface that made its intelligence inspectable.

The design thesis

Transparency is the product.

Instead of hiding Aristotle’s complexity behind automation, we designed the interface around the ways researchers evaluate information themselves.

That meant giving researchers visibility into:

  • what Aristotle is doing
  • why it is doing it
  • how it reached an answer
  • and where they can challenge it

This became the principle behind the Beta experience.

Where we started

Just ask. Aristotle chooses the best model for you.

The first version of Aristotle was intentionally simple.

Every prompt was automatically routed across three models: Explore, Generate, and Verify.

There was no sidebar. No model picker. No visible reasoning layer.

The logic was straightforward:

Early Aristotle explorations: the home screen with Explore, Generate and Verify model chips and match scores, a verification timeline with answer confidence, and the model switcher states

And for simple questions, it worked.

But as Aristotle became capable of doing deeper research, the same simplicity became a limitation.

When simplicity broke

A black box, for people trained to question everything.

Automatic model selection meant researchers couldn’t see why Aristotle chose a particular approach.

When an answer felt wrong, there was little to interrogate.

At the same time, the product was gaining capabilities that couldn’t comfortably live inside the original canvas:

  • Double Check needed room to surface verification.
  • Source previews needed space for evidence.
  • Reasoning traces needed somewhere researchers could inspect them.
  • Different research modes needed to become visible and understandable.

Putting everything directly into the conversation made the canvas increasingly noisy.

Two things had to change.

Researchers needed more control over how Aristotle approached a question.

And Aristotle needed a dedicated space for making its work visible.

Control

How much control should a researcher actually have?

As Aristotle became more capable, we had to rethink an assumption from the original experience: that hiding complexity would make AI easier to use.

Instead, I explored how much of Aristotle’s underlying intelligence researchers should be able to see and control.

The exploration focused on two connected questions:

How should researchers choose how Aristotle approaches a question?

And how should Aristotle expose the work happening behind the answer?

Control · Model control

From automatic routing to researcher intent

Rather than adding more models to a hidden router, we made the choice of approach the product’s core interaction model.

Before

Ask a question → Aristotle decides how to answer it.

After

Ask a question → understand the research intent → choose the appropriate approach.

SparkGenerate hypotheses and explore possibilities.For when you’re exploring what could be true.
SearchFind and synthesize scientific evidence.For when you’re looking for what is already known.
VerifyInterrogate and validate an answer.For when you need to challenge what you have.

A fourth mode, Instant, stays available for quick conversational answers when depth isn’t necessary.

Exposing the choice raised a new problem. Our first explorations surfaced the models directly:

Model picker explorations: automatic Explore, Generate and Verify chips with match scores; three dropdown directions, from plain mode names to modes with descriptions and an External models submenu; and the composer with the picker built in

Model names alone don’t explain when or why someone should use them.

So we reframed the choice. Instead of asking:

Which model do you want?

What are you trying to accomplish?

The model became an implementation detail. The research intent became the interface. Each mode then got its own design pass.

Spark

Spark: from questions to hypotheses

The problem

Scientific research rarely starts with a perfectly formed question.

A researcher might start with an observation, a gap in the literature, or a broad question and need to explore several possibilities before deciding what is worth investigating.

Traditional AI interfaces are optimized for answering questions. Spark was designed for the step before the answer: helping researchers turn an open-ended question into testable hypotheses.

The exploration

The challenge was figuring out how much structure to introduce without turning exploration into a rigid workflow.

We explored how Spark could take a broad research objective, reason across the problem space, and surface multiple potential hypotheses for the researcher to evaluate.

Rather than presenting a single answer, the experience needed to make possibilities the output.

Early Spark exploration: a research question about CRISPR-Cas9 models of drug-induced cardiotoxicity, answered with expandable hypothesis cards tagged novel, low adverse-event risk and testable, with mechanism and a week-one experimental protocol

The final experience

Spark became Aristotle’s deeper exploration mode.

A researcher starts with a question or objective, and Spark works through the problem to generate potential hypotheses, supporting rationale, and directions for further investigation.

The researcher isn’t expected to accept a single conclusion.

Spark turns an initial question into a set of possibilities worth investigating.

Final Spark experience: a valsartan repurposing question answered with scored hypotheses, each opening into a claim, a test, what follows if true, and linked sources

Verify

Verify: from getting an answer to challenging one

The exploration

We explored how Aristotle could expose the different ways it evaluates an answer. Instead of treating verification as something that happens invisibly in the background, the interface could surface individual checks researchers might care about.

The challenge was information density. Researchers needed access to this depth when they wanted it, but showing every check inside the main conversation would make the experience harder to navigate.

Verify exploration: an answer about CRISPR screen confidence broken down by coverage, replicates and statistical approach, beside a verification timeline with answer confidence, source verification, mathematical validation, causal reasoning, Bayesian calibration, key sources and verification coverage

The decision

Make verification available without making it part of every answer.

We moved verification out of the main conversation and into a dedicated layer where researchers could inspect how Aristotle evaluated its response.

Instead of presenting every check at once, we grouped them into individual, inspectable signals. Researchers could quickly see what was checked, then go deeper only when a result raised a question.

This created a balance between visibility and information density: verification was always accessible, without overwhelming the primary research workflow.

The final Verify experience

From a confidence score to something researchers could challenge.

Verify became a space where researchers could inspect the different dimensions behind an answer, from source verification and reasoning consistency to mathematical and factual checks.

Each check gave researchers a way to move beyond a single confidence score and understand where an answer was strong, where uncertainty remained, and what was worth investigating further.

Final Verify experience: the CRISPR screen answer beside a Reasoning panel listing each check with its result, from answer confidence and source verification to mathematical validation, causal reasoning, Bayesian calibration and key sources

The goal wasn’t to tell researchers whether an answer was correct.

It was to give them enough visibility to decide for themselves.

Verify is the model. What it makes possible is the next part of the story.

System Thinking

What happens between the question and the answer?

Choosing the right research mode solved one part of the problem. But giving researchers more control only matters if they can also understand what Aristotle is doing with that control.

System Thinking

Instead of presenting a finished response as a black box, System Thinking gives researchers visibility into the research process behind it.

  1. Question
  2. Scope the problem
  3. Identify relevant concepts
  4. Explore evidence
  5. Connect findings
  6. Form conclusion

Designing for inspectability

The challenge was showing enough of Aristotle’s process to build understanding without exposing every intermediate step.

We designed System Thinking as a progressive view of the research process. Researchers could see how Aristotle moved from the initial question through different stages of investigation, while keeping the final answer as the primary focus.

Each stage could be expanded to reveal the relevant reasoning, evidence, and connections behind it.

How researchers use it

System Thinking gives researchers a way to follow the path, not just read the destination.

They can quickly scan the overall approach, then drill into a specific step when something feels unclear, surprising, or worth challenging.

For example, a researcher might notice that Aristotle identified an unexpected concept during Explore Evidence, open that step, and inspect the sources and connections that led to it.

This makes the AI’s process something researchers can inspect, question, and build on, rather than something they simply have to trust.

System Thinking: a request to generate novel molecules for Lp(a) binding, with each research step checked off in turn, the sources Aristotle read listed beneath it, and a short note on what it will do next

Researchers don’t just see what Aristotle concluded. They can see how it got there.

System Thinking · A place for the work

Giving the work a place to live

We first explored keeping everything within the conversation itself, without a persistent sidebar or traditional navigation. The advantage was simplicity: the researcher could focus entirely on the question and answer.

But as we introduced model controls, verification, sources, and System Thinking, the conversation became responsible for too many jobs at once.

Not another screen.

Not another workflow.

A place for Aristotle’s work to exist alongside the answer.

We introduced a persistent side panel for the work happening behind the answer. The conversation remained the primary place to ask and explore. The sidebar became the place to inspect and understand.

The answerWhat Aristotle is telling you.
The workHow Aristotle got there.
The evidenceWhat supports the answer.
The side panel in two states: Highlights, listing passages the researcher saved with their notes, and Sources, listing the ten papers behind the answer with filters for journal, news and other

Peer Review

What if Aristotle could challenge its own work?

Researchers don’t stop at getting an answer. They question assumptions, look for weaknesses, and ask what might have been missed.

Peer Review extends that behavior to Aristotle itself. Instead of asking researchers to manually identify every weakness in a response, Aristotle can take a second pass on its own work, looking for gaps, unsupported claims, alternative interpretations, and areas that deserve further investigation.

A second set of eyes

The goal wasn’t to create another confidence score. It was to introduce productive disagreement into the research process.

Peer Review asks Aristotle to step outside its original line of reasoning and evaluate the work from a different perspective:

  • What assumptions is the answer making?
  • What evidence is missing?
  • Are there alternative explanations?
  • Which claims deserve closer scrutiny?
  • What would a researcher challenge?

This gives researchers something more useful than “the answer is probably correct.” It gives them places to question.

Designing the interaction

We explored how Peer Review could fit into the research workflow without interrupting it.

Rather than making critique part of every response, we treated it as an intentional second pass. Researchers could ask Aristotle to review the work after an answer had been generated, then inspect the issues it surfaced.

The interaction was designed around a simple loop:

  1. Answer
  2. Review
  3. Challenge
  4. Refine

A researcher could scan the review first, then open individual findings to understand why Aristotle flagged them and return to the original answer with those challenges in context.

What we explored

The exploration centered on one question:

How much should Aristotle challenge the researcher?

Too little, and Peer Review becomes another reassuring summary. Too much, and every answer becomes buried under criticism.

We explored different ways of surfacing critique, from high-level review summaries to individual challenges tied back to specific parts of the response.

The direction that emerged was to make the challenges themselves the interface. Instead of giving researchers another opaque score, Peer Review surfaces specific things worth questioning.

Peer Review: an answer on why pancreatic, ovarian and esophageal cancers evade early detection, beside a Peer Review panel listing the claims it found, each marked verified, unverified or flagged with the reasoning behind it

Peer Review vs. Verify

Verify and Peer Review play different roles in the loop.

Verify

Checks the work. It helps researchers understand whether an answer holds up against the checks Aristotle applies to it.

Peer Review

Challenges the work. It asks what could be wrong, incomplete, or worth reconsidering.

Together, they move Aristotle beyond simply producing an answer:

Verify asks: “Does this hold up?”
Peer Review asks: “What should we challenge?”

And that completes the larger control loop:

  1. Choose the approach
  2. Understand the process
  3. Challenge the result

The goal isn’t to make researchers trust Aristotle more. It’s to give them better reasons to question it.

02 · After Beta

What changed once researchers started using it

Beta helped us validate Aristotle’s core value. But once we put it in front of real researchers and watched them use it, a different set of problems surfaced.

The intelligence was useful, but the product around it was harder to understand. Researchers were unsure where work belonged, what was connected, and how to move between sessions, projects, tools, highlights, and documents. Even “New Session” was interpreted by some as “New Project,” revealing that our information architecture didn’t match their mental model.

New problem

Aristotle worked well as an AI research assistant, but not yet as a coherent research workspace.

Researchers weren’t thinking in isolated chats. They were moving across papers, datasets, notes, reviews, and ongoing lines of inquiry.

Features like Highlights were valuable, but difficult to find and organize, while strong workflows like Peer Review still felt disconnected from the rest of the product.

Where does my work belong, and what is connected to what?

What we learned

The strongest signals weren’t about answer quality. They were about workflow.

The Shape Sessions showed that researchers saw real value in Aristotle, but the strongest signals weren’t just about answer quality. They were about how the product fit into their workflow.

Three themes kept coming up:

01Structure was unclear

Researchers didn’t always understand the relationship between sessions, projects, and tools. “New Session” could even be read as creating a new project.

Make the project the obvious home for work.
02Valuable things were hard to return to

Highlights were valued, but hard to discover and difficult to navigate when organized only by recency.

Give saved evidence structure, not just a timeline.
03Tasks were clearer than models

Peer Review stood out because it mapped to a real research task, while model names and selection often felt ambiguous.

Name workflows by the task, not the model.

The broader signal was clear:

Aristotle needed to organize research around the work researchers were doing, not around the underlying AI system.

What changed

From a collection of AI interactions to a structured research workspace

Instead of allowing chats, highlights, documents, and tools to exist independently, we made Projects the foundation of the product. Every session, document, highlight, and workflow now belongs to a project, giving researchers a persistent place for the context and artifacts around a body of work.

Everything now belongs to a project. Nothing lives independently.

The project became the system of record for research.

Before
Beta project page: a project banner and description, a list of chats that were added to the project, and a separate assets panel of uploaded files
After
After Beta: the sidebar opens on the current project, with Peer Review, Lab Sync, Highlight and In-line definitions as tools inside it, and the project’s sessions listed below

We moved from optional project organization to a persistent project model where every session, document, highlight, and tool inherits the same research context.

We also moved away from asking researchers to understand which model to use. The research showed that model names and modes were often unclear, while task-specific workflows like Peer Review were much easier to understand.

The new direction was simple:

Organize Aristotle around what researchers are trying to do, not around how the AI works.

New features / workflows

That shift led to a set of more explicit research workflows.

These weren’t four unrelated feature additions. They became explicit workflows inside the same project context.

Peer Review: a paper on LRRK2 in Parkinson’s disease open beside Aristotle’s review, which counts claims checked, verified, uncertain and flagged and lists each claim with its supporting evidence
Peer Review

A dedicated surface for reviewing papers and proposals, surfacing weak claims, inconsistencies, and areas that need attention.

Lab Sync: a company preprint open beside Aristotle, which asks how the researcher wants to use the file and then questions how the paper’s methods were validated
Lab Sync

Researchers bring their own data into Aristotle and reason over it within the same project.

Highlights for a Parkinson’s research project: highlights grouped by session with color labels, counts per color, search, sorting and researcher notes
Highlights

Saved passages became an organized research artifact, with color, filters, search, and notes.

In-line Definitions: a term in a paper, “GTPase domain”, opens a definition card with its sources without leaving the document
In-line Definitions

A reading workflow for understanding scientific terminology without leaving the document.

Together, these workflows made Aristotle feel less like a chatbot with modes and more like a workspace built around the different stages of research.

Highlights: from saving information to organizing evidence

Of the four, Highlights has the clearest line from research to design. Researchers liked Highlights but struggled to find and navigate them, especially when sorted only by recency.

Highlights for a Parkinson’s research project: highlights grouped by session with color labels, counts per color, search, sorting and researcher notes
ColorLightweight categorization, without forcing a folder structure.
NotesAdd the researcher’s own interpretation to the evidence.
Project contextKeeps evidence connected to the work it came from.

Impact

A clearer product model

The post-Beta research gave us a much clearer product model. Instead of continuing to add capabilities onto the original chat structure, we reorganized the experience around persistent projects and explicit research tasks.

The result was a product with a clearer mental model:

One project, one research context, multiple connected workflows.

It also gave us a stronger foundation for future features. New tools, documents, notes, and research artifacts could now inherit the same project context instead of becoming another disconnected surface.

Beta taught us how researchers wanted Aristotle to answer.

After Beta taught us how Aristotle needed to fit into their work.

03 · Reflection

What I learned designing AI for researchers

Designing Aristotle changed how I think about AI products.

At first, I spent a lot of time thinking about the quality of the model: how much reasoning to expose, how to distinguish different modes, and how to help users choose the right one.

But the research showed me that users rarely want to think about the system at that level.

Researchers cared about whether Aristotle could support the task in front of them: review a paper, understand a term, work with a dataset, save evidence, or challenge a claim. When we exposed too much of the underlying AI architecture, we often made the product harder to understand. Model selection itself became a source of uncertainty, and visible reasoning could compete with the answer rather than build confidence.

I also learned that trust is not the same as transparency.

Researchers did want to understand what Aristotle was doing, but more information was not always more useful. What mattered was showing the right evidence, uncertainty, sources, and reasoning at the moment it helped them make a judgment.

The biggest shift in my thinking was from designing an AI interface to designing a research system around AI.

The model was only one part of the experience. The harder design problem was creating the structure around it: where work lives, how context persists, how evidence is organized, and how researchers move from asking a question to building knowledge over time.