This paper shows that machine learning (ML) retrosynthesis models such as 3N-MCTS and AiZynthFinder incorporate non-epistemic values following the traditional inductive risk problem in philosophy of science. We argue that inductive risk in these cases is distributed, since it arises from the choice of model to construction, the data templates and to implementation. We connect this to traditional inductive risk but due to its proliferation throughout the process, call it distributed inductive ris…
Read moreThis paper shows that machine learning (ML) retrosynthesis models such as 3N-MCTS and AiZynthFinder incorporate non-epistemic values following the traditional inductive risk problem in philosophy of science. We argue that inductive risk in these cases is distributed, since it arises from the choice of model to construction, the data templates and to implementation. We connect this to traditional inductive risk but due to its proliferation throughout the process, call it distributed inductive risk. Second, we identify a novel source of inductive risk generated by algorithmic vagueness. Non-deterministic ML models cannot interpret vague terms and therefore must define notions like ‘chemical similarity’ into precise, operational categories. This involves moving beyond purely technical adaptations and ultimately points to a gap that is filled by appealing to value judgements. Unlike traditional inductive risk, which operates through explicit trade-offs between error types, algorithmic vagueness produces mandatory value commitments at the structural level.