I do not think this is about the system being not net positive, but more about what could be different. There are much more actionable feedback we can give about reproducibility. We can also flag wrong claims in papers with much better precision and recall. I do not think I could run this even a year ago. The complexity and diversity of different papers’ code repos are huge.
This may be a trust example (but hard to judge without author response or code release). These kind of issues were always there. Peer review is broken, but these issues of reproduction are not new, and they didn't stand in the way of the system being an overall (huge) net positive when it worked (not without flaws). My sense is that you could have done this study 15 years ago with similar results