Vision
What is a programmable cloud laboratory, and what stands in the way of making such facilities ubiquitous?
A programmable cloud laboratory starts with an experimental goal, plans the work, runs the experiment on real instruments, reads the results, and decides what to do next — without a scientist at the bench. The computer-enabled instruments for creating programmable cloud laboratories already exist — our node alone has more than 740 of them, spanning analytical and synthetic chemistry, to biomedicine and ’omics, to materials science. But the general-purpose specification layer for enabling these instruments in a standardized, platform-independent manner still needs to be built.
The bottleneck is method transfer
Laboratory methods are written for human interpretation: critical details are often left out, results are difficult to reproduce, and procedures are hard to scale across institutions. Through fifteen years of operation, Emerald Cloud Lab has found that the hardest part of running an automated experiment is writing the specification in the first place. A formal, complete protocol is exceedingly difficult to write, even with experience and training. The usual failure is under-specification: researchers are not accustomed to thinking through the low-level details that are normally left unsaid — but that are essential for the operation of automated systems.
Laboratory automation today is built on rigid specification languages that are tied to particular instrument models and software packages. The language used at Emerald Cloud Lab, Symbolic Lab Language (SLL), is detailed, specific to its facility, and not easily adapted elsewhere. Translating written methods into SLL requires trained specialists. What is missing is an intermediary layer between the way investigators think about their experiments and laboratory-specific languages such as SLL.
From a written page to a running experiment
Our pipeline has five stages. Select a stage to see what happens there.
A written protocol
A textual document describing a laboratory experiment, written in prose for a trained human reader. Critical details — timings, temperatures, tolerances — are often left unsaid, so different laboratories may get different results from the same page.
ORE — an agent that resolves ambiguity
The Operational Rendering Engine reads the method and finds what it leaves ambiguous. It resolves each gap through question-and-answer dialogue with the researcher — or, when no one is available to ask, through small AI-designed experiments.
AMBER — a shared language for experiments
The refined method becomes an AMBER (Abstract Model for Bridging Experiment Representations) workflow: a typed, machine-actionable graph of operations, resources, parameters, and provenance. It is an open standard, tied to no single facility or vendor.
TROVE — a translator into node languages
TROVE converts the AMBER representation into each node’s own execution language — beginning with Emerald Cloud Lab’s Symbolic Lab Language, and extending to new node languages as the Test Bed network grows.
Execution — and the closed loop
Programmable instruments run the experiment at a cloud-lab node. Results feed back through the pipeline — adjusting parameters, updating constraints, and proposing improved protocol variants. That closed loop is what lets a method improve run over run — and every cycle of it publishes its outputs as FAIR objects: workflows, data, and provenance that are findable, accessible, interoperable, and reusable by any node in the network.
Three technical objectives
A representation for experiments
The Abstract Model for Bridging Experiment Representations is a typed, machine-actionable schema language that captures laboratory operations, resources, parameters, and provenance in an implementation-agnostic form. Knowledge representation is implemented using semantic technology, with each workflow modeled as a directed graph of control-flow and data-flow dependencies.
An agent that resolves ambiguity
The Operational Rendering Engine inspects written methods and resolves what they leave unsaid, through user dialogue where a user is available and through small, AI-designed experiments where one is not. Where a protocol says “allow to equilibrate,” the agent can systematically vary the underspecified parameter until results agree with known standards.
A translator into node languages
TROVE converts laboratory-independent AMBER representations into a PCL node’s particular execution language. Our work will begin with the SLL language used at Emerald Cloud Lab, and it will have the potential to extend to any node-dependent language as the NSF PCL Test Bed network grows.
Across all three components, benchmarking, provenance tracking, and reliability assessment are built in from the start, so that autonomous experimentation stays safe, auditable, and reproducible.
The science driver
Since the 19th century, the U.S. Pharmacopeia (USP) has been publishing textual monographs that describe analytical laboratory procedures for assessing the purity and quality of drugs and other compounds. Although the procedures described in these monographs are intended to serve as standard protocols, different laboratories — and different technicians within one laboratory — may execute them differently, and small differences in composition, solvent, or temperature may alter the result. The interpretation of USP monographs and the execution of the associated laboratory procedures are well-constrained problems that emphasize reproducibility and precision. The conversion of USP monographs into actionable specifications for a programmable cloud lab is a great first test of our approach — and one with direct consequences beyond the lab, demonstrating scalable pharmaceutical quality testing in support of resilient and secure domestic manufacturing. In the later years of our project, we will extend our work to address broader science drivers.
A network, not a node
AMBER is designed to be independent of any single facility: TROVE translates AMBER specifications into each node’s internal language, beginning with Emerald Cloud Lab’s SLL. GEMSTONE hopes to serve as the interoperability hub for the NSF programmable cloud laboratory network. Working with the other PCL nodes on APIs and data standards, our team hopes to spearhead the language-design effort that will allow nodes to share, validate, and execute experiments through common, machine-actionable specifications.
15 translations, built and maintained by hand — for just the six nodes shown. Every pair of execution languages needs its own translation, so the work grows with the square of the network: ten labs need 45, twenty need 190.
6 translations — one per node, however many nodes there are. Each lab connects once, to the shared standard, and can exchange, validate, and execute every other lab’s experiments. Twenty labs need 20 translations, not 190.
Illustrative — Emerald Cloud Lab is the network’s first node. The other nodes stand for the kinds of labs that can join, not specific partners, and the network is open-ended.
AMBER is meant to do for laboratory automation what the STEP language did for computer-aided design and manufacturing in the last century: by creating a shared interchange layer, AMBER will enable the PCL community to avoid becoming siloed, enabling the widespread sharing of specification languages, of protocol-authoring tools, and of whole laboratory facilities.
Why it matters
The development of shared standards and AI-based translation approaches will lead to faster discovery, more reproducible results, and broader access to advanced experimental capability — for universities, startups, nonprofits, government laboratories, and small businesses that have never before had a route into automated experimentation.
Along the way, the project will train a workforce fluent in AI, biotechnology, and advanced manufacturing — an investment in U.S. competitiveness, pharmaceutical supply-chain resilience, economic growth, and national security.