Looking at this thread, I can see that a lot of technical people have ambivalent to negative feelings towards AI, but with each new generation, I become more and more convinced that they're missing out on something interesting.
It is indeed true that all models are, at their core, predictors of what occurs next in a sequence. But I think it's worth exploring the implication of what that means. Because when fed tiny pieces of information for a few tasks at a small scale, this results in something that sorta, kinda works. Or, works surprisingly well.
But when scaled... When the amount of information starts approaching the sum of all human knowledge, the tasks start approaching all useful applications of that human knowledge, and the fidelity of the predictor approaches incomprehensible sizes, the starts encodes / becomes (I'd argue it becomes) something that can model all human knowledge.
It feels wrong to say that, but let me explain, what is the best way to predict the behavior of a ball constrained in two directions that bounces with initial vertical velocity v(y) (y is up / down axis) and horizontal velocity v(x) (x is side by side in 1d) ?
If we purely look at it via a graph, it's by modelling the function of acceleration under earth's gravity.
If only a few points are given to you for this and you can't make something really sophisticated, then you'll make something that's rough that kinda sorta works and then call it a day.
But... if the number of points keeps increasing in number, precision and accuracy as well as the number of examples (assumed that data about air pressure, velocity and all other factors is included alongside these points), the fidelity with which you can replay / tweak the function keeps improving, and the number of times you can iterate keeps increasing, you'll eventually create a function that models that process so well that it intrinsically contains a good enough model of the deformation of the ball (provided the dataset contains information about elasticity of the ball's material, its dimensions and mass etc..), the nearly negligible (under normal conditions) effects of the ambient environment (provided there's diversity in the number of environments supplied), the oblateness of the Earth and minute changes in the gravitational field (the length of a seconds pendulum varies depending on where the experiment happens. It's presumed that all of the prior set of experiments were repeated across the Earth and the subtle, but real deviations were faithfully recorded)... and so much more.
A machine trained on the above with a large number of parameters, measures to prevent "laziness" and enough reps for high fidelity across a large enough dataset would start to approach a simulation of the ball falling. Because to predict what happens next in the sequence, you must model what's occurring in the sequence.
Now imagine doing that for other tangible and intangible things in this world. For all of human knowledge across all fields of endeavor. All experiences. No matter how noble, ignoble, notable or ignorable. But putting all of it into the soup that's this machine. Then at larger and larger scales, you eventually start encountering "good enough" models (in modelling the falling ball sense) for even the most hard to quantify / qualify things like grief and joy. At some point, by simply trying to predict what it has been taught ought to be the next part of the sequence in say... human interaction, it starts to make a model of something that hews ever closer to a full fidelity theory of mind.
Is there evidence for this? Kind of, yes. There are early indications that as machines are trained for an ever larger number of tasks at larger and larger scales, their internal representations converge. It's called the Platonic Representation Hypothesis. Overview and paper here, https://phillipi.github.io/prh/
It is my opinion that these machines are displaying a new form of intelligence that human beings haven't quite encountered before. They are the sum of all human knowledge made manifest and given voice by processes that nudge (bit-by-bit) what kind of step it ought to predict for the next part of whatever sequence it displays.
In my mind this means that, of course, these models can create new knowledge. This strains the analogy, but with the sum of all human mathematics within them, they can "reason" via the act of predicting what ought to come next.
Of course, these machines are "surprisingly" good at a lot of things the larger they get, because what the labs have created here is a rough version of humanity's collective knowledge given form and the ability to say hello.
I suspect that the current generation isn't close to the "true frontier" of what these machines could be. They are nowhere close to the sum of all human knowledge and endeavor. They are quite a way there, but they haven't yet achieved true completeness for domains where the data isn't so public.
I think it's the most exciting scientific and technological breakthrough of my lifetime. And I can't wait for us to get close to the true frontier of all domains.