They don't need to back it up. It is up to the person making the original claim to back it up.
There's half a century of research and practice into this topic. Anyone in the industry _should_ know about this. Anyone who doesn't can find it on a search engine within 30 seconds.
No, it's a well known problem. The old lines of code joke still applies. You can introduce metrics, but those will just be gamed because software is a multidimensional optimization problem with complex trade offs. So, you can't translate that to a single measurement and call it productivity.
Before which you have to figure out how to measure quality. There are countless ways of course, but choosing your metrics is subjective and pretty much arbitrary.
Also, since many tasks are literally only done once, "task" is itself a variable, not a fixed data point. As for "time", that can be quite hard to measure, since unless it's a very menial task it'll span several hours, or days even, and it's quite elusive to total what really constitutes real time spent on a given task.
I believe productivity/efficiency in software engineering is something that is felt (i.e you can measure some of its effects but many are a matter of perception, not definitive metrics. like pain.) rather than measured (i.e you can synthesize a definitive metric from a set of known, well-defined data).
Let's imagine that I wrote some test cases. To me, If my piece of code passes those tests, I will consider it of being of X quality. I won't care about any other quality than this X one.
If I write/build/test two different pieces of code, both passing my tests, and happens that I took less time to write the second one of them, then I would consider I was more productive writing the latest.
I think your example shows exactly why measuring software productivity is so problematic. With your scenario, probably the easiest way to "improve productivity" would be to not write any documentation comments, heck it would probably be faster to just stick it all in one giant method.
I think many of us have known software developers who are incredibly fast, but who write "rat's nest" code that's impossible for anyone else to decipher, or which is very difficult to modify when new requirements come along. Quantifying those attributes is incredibly difficult, but as most time in software is actually spent on maintenance, it usually ends up being more important than a simpler assessment of initial output.
One can assume writing documentation and comments is covered by both approaches. Still, the fastest scenario would be the most productive one.
I agree with your view regarding maintenance cost, but I don't think that's what's at stake here. We can go further and try to compare the amount of time needed to fix rat's nest with the amount of time needed to write the software all over again. Covering the rat's nest is in many occasions faster (on which I'm considering more productive) than rewriting.
You're never going to get around the fact that most of these things are qualitative assumptions. Is your documentation good? There's really no completely objective way to measure maintainability, because maintainability is often about how easy it is to handle unknown future change requests.
Years ago "function points" were bandied about as a truly objective measure of software output and value. It turned out to be trivially game-able and added 0 value to software development estimation.
Where is your metric? By how much will you be declared more productive? Is it just time to achieve a result? Because that is obviously wrong. How are you accounting for quality, or even measuring that X? Do you account for readability? maintenance? documentation? factoring? over-engineering?
There are so many visible as well as hidden factors that attempting to measure any (let alone all) of that is a pipe dream.
Also, by writing the second one you are not accounting for second version effects since it is a reimplementation of something you know, and if someone else develops it for you then you just changed a major variable.
I get that you said "oversimplified", but this is just like saying "thought experiment", yet as soon as you confront that to the real world, it just doesn't realistically fly.
Ok, so now you "just" need to factor in long-term maintenance.
Besides, how often do you write code twice just to see which way is better? What you have is a comparative measure that is completely useless for measuring new code, i.e. code which implements a new feature.
Even it you just hardcode the tests cases in a giant lookup table?
Tests can prove the presence of bugs, but paying doesn't mean the software is bug free.
I don't know if I would agree that there are countless ways to measure quality. I think the standard one is: the ability of the software to meet its design specification. Depending on your spec this could be more or less objective.
I would say that's a good baseline, but I would extend that to "the ability of the software to meet its design specification well", where real quality lies in the more subjective subtleties between minimally functional and genuinely happy users.
Would you like to back your statement up?