It’s an interesting and important question and the answers are not straightforward:
The cost functions for neural networks are not convex. There is no guarantee that you’ll ever find a global minimum, but that may not be desirable in any case: most likely it would represent extreme overfitting on the training data. If you choose your gradient descent algorithm and parameters correctly, you can usually find one of the very many local minima that give a reasonable solution.
Here’s a thread which discusses this in more detail.