In theory, the solution in situations like this is to continue the test over a longer time. Typically with ‘holdbacks’ – subsets of users who don’t get a feature for a long time. This is easier if you have an app that everyone uses because with a website, it’s harder to reliably find holdbacks who are also a representative sample (eg it won’t work so well to hold back everyone using the site in a certain language as those people may be statistically different from the general population in other ways).
There are still a bunch of problems – higher maintenance burden, harder to iterate on a site quickly. Though I think you identify what I would consider the bigger problem which is that they cause political difficulties as a holdback can only really turn around and say that the positive impact people claimed wasn’t really borne out in the long term. So even at places that do holdbacks, the results may be silenced or ignored. If a holdback shows something continuing to work, that’s hard to get excitement about even though I think one should expect many of these a/b test results to not have long lasting effects.
Am I the only one who thinks THAT is a dark pattern? It would at a minimum, confuse me, if I had a different UI / feature set than my friend(s) - confused for being unaware why my experience looks different, and you better bet that 90%+ of users do not know what an A/B test is. At worst, I'm angry because they are getting a better experience and I randomly got shafted with no ability to upgrade.
That's the same kind of thinking that would criticize pharmaceutical companies for giving placebos. They do it because they don't know with confidence, and they want to. They are collecting evidence.
To get into a medical trial you a) know it's a trial b) get informed that you might get a placebo c) consent to participate in the trial.
If a pharmaceutical company was discovered filling 30% of pill bottles they sent to pharmacies with placebos, they'd be sued unto the eighth generation.
The types of effects you want to measure would take months or years to show and are the combination of many different small decisions. Teams need to apply common sense thinking, empirical data, and the willingness to wrestle with uncertainty. All of that is hard, blindly following a/b test results relieves people of that cognitive burden
There are still a bunch of problems – higher maintenance burden, harder to iterate on a site quickly. Though I think you identify what I would consider the bigger problem which is that they cause political difficulties as a holdback can only really turn around and say that the positive impact people claimed wasn’t really borne out in the long term. So even at places that do holdbacks, the results may be silenced or ignored. If a holdback shows something continuing to work, that’s hard to get excitement about even though I think one should expect many of these a/b test results to not have long lasting effects.