Of course it is and clock references are an extremely poor analogy IMO.
If real differences, eg something you've measured and know to exist aren't identifiable on a blind test, that tells you that the difference while real isn't audible, probably doesn't matter and you can ignore it.
I can (and have) shown for example that small adjustments to EQ of a db or two are clearly and repeatedly identified on a blind test. Drop the variation to 0.25db and it becomes harder, and often impossible to reliably detect.
That simply tells me that we've established where a real difference either matters or doesn't.
Do the same thing sighted and the difference - no matter how small - is always heard becasue the listeners are fully aware of what you've done.
Now, neither test is perfect but I know which I'd being more inclined to rely on.
Trouble is, the test you have described is scientifically pointless. It's laden with potential biases and errors. If you ran this test under lab conditions with certain material, you'd find panels fail to find differences even with 6dB changes in an EQ curve, but with other material, 3dB would be the threshold. Anything less than that is fantasy land, from a scientific perspective.
Play different EQ curves in one order and you get one set of results, play them again in another order and you'll get a completely different set of results. Play more than A-B and you bias listeners toward the middle presentations.
So your 'real difference' is no such thing. It can confidently be ignored too, even though it appeared under blind conditions, because those conditions weren't blind enough. In other words, it's no more or less robust than someone being shown what they are listening to before they listen to it.
Worse, given the huge range of variation in tests depending on how you stage the test, how can you be sure that if something fails to pass your test that it isn't that your test is throwing out a false negative? You are making a huge assumption on the validity of your test, based purely on the test results. That way lies all manner of random results.
From my experiments, no one methodology is overarching. The only way to test things is to do use as broad a spread of tests as possible and weight their scores according to their reliability.
But this is a topic that is probably best broken out of this thread, as it has no relevance to the OP.
Last edited by a moderator: