Human-Driven Scientific AI
Updated · 4 episodes · 1 show · 4 source notes
Definition
Human-driven scientific AI uses models to prepare data, detect patterns or suggest experiments while domain researchers choose questions, inspect outputs, verify claims and control physical risk.
Current Synthesis
Across four interviews from Data Science With Sam, the binding constraint changes by domain: molecular-data representation, sparse spaceflight events, interpretability of EEG categories, and the quality and safety of laboratory records. Sam (Data Science With Sam)’s “human-driven” formulation is an editorial stance of the speakers, not a comparative demonstration that all human-in-the-loop designs work.
Key Claims
- Scientific usefulness begins with fit between data representation and a domain question, not model size alone.
- Incomplete records, missing failed experiments and disciplinary translation gaps limit model suggestions before any physical experiment.
- Candidate synthesis routes and experimental decisions require biological or chemical interpretation, reproducibility and safety review.
- In spaceflight, sparse one-off events favor bounded imagery-review tasks over unconstrained automation.
- Brain-signal classification and assistive ambitions require replication and user validation; predicting an object category is not reading thoughts.
Evidence
- Data to interpretation: EP 8: Implementation of AI in scientific research describes Lucas Simon’s Baylor Therapeutic Innovation Center workflow: Sequencing Data Pipeline raw reads become a Gene Expression Matrix of roughly 20,000 genes; Computational Biology interprets it downstream of Bioinformatics. Single-Cell RNA Sequencing changes bulk experiments of hundreds of samples to tens of thousands or up to about a million cells. A Single-Cell Autoencoder Representation is meaningful only if its clusters map to cell types or testable biology, not merely a neat visualization.
- Records and experiment selection: Data, AI, and Scientific Research: A Coffee Chat reports Effie (Data Science With Sam)’s concern about Experimental Science Data Quality, SOP deviations, blinding, and the Bioinformatics Domain Gap; the stained-tissue mutant/wild-type clustering example had a blinded developer. Mossam (Data Science With Sam) discusses Retrosynthesis AI and Negative Results As Scientific Data: published successes omit many failed reactions; fluorine-18 or carbon-11 Radiochemistry Imaging Tracers constrain timing of final labeling, and proposed routes still need chemist review. The episode mentions blind comparisons of routes without an independent accuracy measure. Blood-Brain Barrier Prediction uses lipophilicity, pKa and polar surface area as candidate filters; AI Experiment Documentation by cameras remains a proposal.
- Sparse and abundant space data: EP 4: A.I. talk with a Rocket Scientist from NASA has Kofi Browning explain Spaceflight AI Dataset Scarcity at NASA but contrast it with International Space Station imagery: Space Imagery AI can triage uneventful footage. EVA Glove Inspection AI assists photographed glove review before spacewalks, including an engineer’s Microsoft collaboration; mission control remains responsible. Sam’s Artemis lunar-rock classification by texture and curvature is a proposed example, not proven flight deployment; AI Model Bias Governance also matters when inputs are omitted.
- Assistive validation: EP 6: Data Science & AI Talk describes Paulina Nemkova’s EEG Brain Reading work at University of North Texas, replicating/extending related Stanford work and classifying thought-about object categories. Locked-In Syndrome Assistive Communication motivates the research, but the episode does not establish a deployed communication device; Research Replication Integrity and AI Research Literature Currency constrain that inference.
Counterevidence & Qualifications
All four notes are episodes of one show rather than independent trial evidence. Scientific Discovery Automation may be useful for routine analysis, but neither a route proposed by software nor a visually plausible cluster proves a result. The speakers do not treat scientific creativity or novel reaction design as reducible to pattern completion. Mossam (Data Science With Sam) raises unsupervised radioactive reactions as a prospective hazard, not a documented trial. Effie (Data Science With Sam)’s concern about unknown biology and Research Taste limit confident extrapolation; Problem Definition In Research cannot be outsourced simply by collecting more data.
What Changed
- Recast replacement rhetoric as four distinct verification boundaries with explicit proposed-versus-observed status.
- Distinguished data-preparation, model representation, experiment selection and physical safety.
Related Concepts
- AI For Science - wider field in which these bounded researcher-led workflows sit.
- Human Judgment Under AI - covers responsibility for interpreting model outputs before action.
- Domain Expert Alignment - explains why bioinformaticians and biologists must reconcile representations and questions.
- AI Verification - turns a model suggestion into a claim tested against experimental evidence.
- AI Research Literature Currency - changing prior work constrains the EEG project’s novelty and replication claims.
- Biomedical Deep Learning - single-cell-scale modeling must retain biological interpretability.
Sources
4 source notes across 1 show
- EP 8: Implementation of AI in scientific research Data Science With Sam
- EP 6: Data Science & AI Talk Data Science With Sam
- EP 4: A.I. talk with a Rocket Scientist from NASA Data Science With Sam
- Data, AI, and Scientific Research: A Coffee Chat Data Science With Sam