The existing philosophical debates on the use of machine learning (ML) in science have organised around the issue of model opacity. What these debates often leave unanalysed is the source of a model's epistemic validity: the overarching ML pipeline that includes dataset collection, model training, and model evaluation. In order to carry out this analysis, I utilise the activity-based framework developed by Hasok Chang, in which the pipeline can be viewed as an integrated activity where each step…
Read moreThe existing philosophical debates on the use of machine learning (ML) in science have organised around the issue of model opacity. What these debates often leave unanalysed is the source of a model's epistemic validity: the overarching ML pipeline that includes dataset collection, model training, and model evaluation. In order to carry out this analysis, I utilise the activity-based framework developed by Hasok Chang, in which the pipeline can be viewed as an integrated activity where each step has its own success criterion. In addition to the steps' individual success, the overall success of the pipeline depends on their coherent coordination. In the course of this analysis, I show that the epistemic validity of the pipeline crucially depends on the operational provenance of the data, which allows the agents to assess the data's adequacy for the pipeline's aim through two requirements (Discrimination and Traceability). The upshot is that opacity is decomposed into a number of separately diagnosable issues, of which the model complexity is only one. Furthermore, link uncertainty is revealed to be a property of dataset collection, and not of the model.