TL;DR GPT-6 Astra has started another familiar AI conversation. The model is more capable, Jensen Huang said on X that "AGI has arrived," and social feeds quickly moved between excitement, curiosity, and fear about what comes next. I understand why the reactions are strong. AI is changing quickly, and some of the things these models can do now would have sounded surprising not long ago. But I have also noticed how quickly "AI is becoming much more capable" can turn into "developers will not be needed" or "CSE is finished." That feels like a much bigger conclusion. For me, Astra raises another question that is easier to miss: How do we actually know an AI model is getting better? We usually look at benchmarks, and they are useful. But benchmarks can age too. Some become saturated, some use proxies for capabilities that are difficult to measure directly, ground truth can be complicated, and evaluation methods can change. A benchmark score also does not automatically tell us how useful a model will be for our own work. Astra's system card gives several examples of this, including older evaluations becoming saturated or being considered for retirement, newer evaluations using more realistic or experimental data, and evaluation methods changing over time. That matters for developers too. If AI makes writing code cheaper and faster, software engineering does not suddenly disappear. Requirements, architecture, system design, security, verification, tradeoffs, maintenance, and deciding what should actually be built still matter. AI will change the work. I think that part is pretty clear. The more useful question is what we do with that change, and whether we are measuring AI well enough to understand what is actually happening.
Another AI launch, and the conversation gets loud very quickly A benchmark gives us a number. What does the number really tell us? Every benchmark has a lifespan 92% according to what? The number can change because the test changed One answer is starting to tell us less So what should we look at when we see a benchmark? What does all of this mean for developers? We can take AI seriously without making every headline the final answer The models are changing. Our way of looking at them has to change too. I Would Love to Hear From You References 🤝 Let's Stay Connected
Whenever a major AI model arrives, there seems to be a familiar sequence. The model launches, people try it, benchmark results appear, comparisons start, and then the bigger questions arrive.
With Astra, NVIDIA CEO Jensen Huang also said on X that "AGI has arrived" while congratulating OpenAI on the launch.
As you would expect, people started interpreting that statement in very different ways. Some saw it as a major milestone, some questioned whether Astra actually meets their definition of AGI, and others immediately started thinking about what this could mean for jobs and software development.
From what I have been seeing in my own feed, there is also a lot of fear around where developers fit into all of this.
I am not saying fear is wrong. There are real risks with increasingly capable AI, and I think those risks deserve serious attention.
But I think there is a big difference between saying "this technology is changing quickly" and saying "developers are no longer needed" or "there is no need for CSE anymore."
Software development will change. There is very little doubt about that. But when producing code becomes easier, other parts of the work can become more important too.
Understanding the problem, designing the system, making architecture decisions, working through requirements, thinking about security and reliability, verifying what was produced, understanding tradeoffs, and knowing when an AI-generated solution is simply the wrong solution are all still part of building software.
None of those become less important just because writing the code becomes faster.
So while everyone is asking how capable the new model is, I started wondering about something slightly different:
There is something very convenient about benchmarks. You see Model A at 82% and Model B at 88%, and it suddenly feels like we have a clear answer about which one is better.
And honestly, benchmarks are useful. Without them, comparing AI systems would be much harder. Having a common test gives us a starting point, and that is valuable.
The part that is easy to forget is that a benchmark is a measurement tool. It gives us evidence about a model's capability, but it is not the capability itself.
Think about a thermometer. It tells you the temperature, but the thermometer is not the temperature. In the same way, a benchmark can tell us something about what a model can do, but it cannot capture everything the model is capable of doing in the real world.
And this becomes more important as models improve, because the test itself can eventually become less informative.
A benchmark that once showed a clear difference between models might eventually have almost every strong model scoring near the top. At that point, the number is still a number, but it may not tell us as much as it used to.
