You must have misunderstood the statement: «In order to properly engineer things, we must know how they work»: of course they may work anyway, of course there may be emergent properties (as was explicitly written), but the deontic part of knowledge augmentation, the scientific part, is missing.
AI is full of non deterministic devices like genetic algorithms, but there exists a problem of transparency which is paramount in the discipline. Automatically produced solution S works: why does it seem to work? Where would it fail? We must understand that.
No, I think you misunderstood me. We engineer the model-training algorithm and then we USE the model. The model is a mess of random gobbledegook that happens to be the best optimizer for the training algorithm's reward function.
This is still science, in the same way it's still science when we breed cows to produce more milk without understanding the full bovine genome.
If something isn't right with the model we pour extra effort into our science - the training algorithm - and train a new one.
At no point does anything ever require peering into the slurry of random digits that make up the model itself, any more than using a computer requires understanding its DRAM training values computed at run time, or any more than pouring water requires understanding the laminar flow along the pitcher's surface (an unsolved math problem!).
> in the same way it's still science when we breed cows to produce more milk
If you cross cows and the amount or quality of milk varies, you have a question in the queue: why did it happen. The question is scientifically duly in abstract terms, and is practically, technically more duly if we benefit from a theory of cows crossing to achieve targeted milk yield.
> require peering into the slurry of random digits that make up the model itself
It is required by the duty of understanding the engine. When we understand engines, we can build better engines. And when we understand things, already our general knowledge is richer.
AI is full of non deterministic devices like genetic algorithms, but there exists a problem of transparency which is paramount in the discipline. Automatically produced solution S works: why does it seem to work? Where would it fail? We must understand that.