Vision

What is a programmable cloud laboratory, and what stands in the way of making such facilities ubiquitous?

A programmable cloud laboratory starts with an experimental goal, plans the work, runs the experiment on real instruments, reads the results, and decides what to do next — without a scientist at the bench. The computer-enabled instruments for creating programmable cloud laboratories already exist — our node alone has more than 740 of them, spanning analytical and synthetic chemistry, to biomedicine and ’omics, to materials science. But the general-purpose specification layer for enabling these instruments in a standardized, platform-independent manner still needs to be built.

The bottleneck is method transfer

Laboratory methods are written for human interpretation: critical details are often left out, results are difficult to reproduce, and procedures are hard to scale across institutions. Through fifteen years of operation, Emerald Cloud Lab has found that the hardest part of running an automated experiment is writing the specification in the first place. A formal, complete protocol is exceedingly difficult to write, even with experience and training. The usual failure is under-specification: researchers are not accustomed to thinking through the low-level details that are normally left unsaid — but that are essential for the operation of automated systems.

Laboratory automation today is built on rigid specification languages that are tied to particular instrument models and software packages. The language used at Emerald Cloud Lab, Symbolic Lab Language (SLL), is detailed, specific to its facility, and not easily adapted elsewhere. Translating written methods into SLL requires trained specialists. What is missing is an intermediary layer between the way investigators think about their experiments and laboratory-specific languages such as SLL.

From a written page to a running experiment

Our pipeline has five stages. Select a stage to see what happens there.

A written protocol

A textual document describing a laboratory experiment, written in prose for a trained human reader. Critical details — timings, temperatures, tolerances — are often left unsaid, so different laboratories may get different results from the same page.

Three technical objectives

AMBER

A representation for experiments

The Abstract Model for Bridging Experiment Representations is a typed, machine-actionable schema language that captures laboratory operations, resources, parameters, and provenance in an implementation-agnostic form. Knowledge representation is implemented using semantic technology, with each workflow modeled as a directed graph of control-flow and data-flow dependencies.

ORE

An agent that resolves ambiguity

The Operational Rendering Engine inspects written methods and resolves what they leave unsaid, through user dialogue where a user is available and through small, AI-designed experiments where one is not. Where a protocol says “allow to equilibrate,” the agent can systematically vary the underspecified parameter until results agree with known standards.

TROVE

A translator into node languages

TROVE converts laboratory-independent AMBER representations into a PCL node’s particular execution language. Our work will begin with the SLL language used at Emerald Cloud Lab, and it will have the potential to extend to any node-dependent language as the NSF PCL Test Bed network grows.

Across all three components, benchmarking, provenance tracking, and reliability assessment are built in from the start, so that autonomous experimentation stays safe, auditable, and reproducible.

Inside a glovebox at Emerald Cloud Lab: an analytical balance reading 0.00000 g, flanked by two mounted cameras that record every measurement, with QR-code labels on the equipment and work surfaces.
Provenance capture at Emerald Cloud Lab: cameras mounted inside a glovebox record every measurement on the analytical balance, and QR codes identify each piece of equipment to the system.

The science driver

Since the 19th century, the U.S. Pharmacopeia (USP) has been publishing textual monographs that describe analytical laboratory procedures for assessing the purity and quality of drugs and other compounds. Although the procedures described in these monographs are intended to serve as standard protocols, different laboratories — and different technicians within one laboratory — may execute them differently, and small differences in composition, solvent, or temperature may alter the result. The interpretation of USP monographs and the execution of the associated laboratory procedures are well-constrained problems that emphasize reproducibility and precision. The conversion of USP monographs into actionable specifications for a programmable cloud lab is a great first test of our approach — and one with direct consequences beyond the lab, demonstrating scalable pharmaceutical quality testing in support of resilient and secure domestic manufacturing. In the later years of our project, we will extend our work to address broader science drivers.

A network, not a node

AMBER is designed to be independent of any single facility: TROVE translates AMBER specifications into each node’s internal language, beginning with Emerald Cloud Lab’s SLL. GEMSTONE hopes to serve as the interoperability hub for the NSF programmable cloud laboratory network. Working with the other PCL nodes on APIs and data standards, our team hopes to spearhead the language-design effort that will allow nodes to share, validate, and execute experiments through common, machine-actionable specifications.

15 translations, built and maintained by hand — for just the six nodes shown. Every pair of execution languages needs its own translation, so the work grows with the square of the network: ten labs need 45, twenty need 190.

6 translations — one per node, however many nodes there are. Each lab connects once, to the shared standard, and can exchange, validate, and execute every other lab’s experiments. Twenty labs need 20 translations, not 190.

Illustrative — Emerald Cloud Lab is the network’s first node. The other nodes stand for the kinds of labs that can join, not specific partners, and the network is open-ended.

AMBER is meant to do for laboratory automation what the STEP language did for computer-aided design and manufacturing in the last century: by creating a shared interchange layer, AMBER will enable the PCL community to avoid becoming siloed, enabling the widespread sharing of specification languages, of protocol-authoring tools, and of whole laboratory facilities.

Why it matters

The development of shared standards and AI-based translation approaches will lead to faster discovery, more reproducible results, and broader access to advanced experimental capability — for universities, startups, nonprofits, government laboratories, and small businesses that have never before had a route into automated experimentation.

Along the way, the project will train a workforce fluent in AI, biotechnology, and advanced manufacturing — an investment in U.S. competitiveness, pharmaceutical supply-chain resilience, economic growth, and national security.