Settings

Theme

Dear Software Makers

blog.jim-nielsen.com

116 points by speckx · 87 comments

Reader

21 threads
beloch

"There’s this sort of unspoken rule of community on YouTube, that like we’re all having this same experience together — we’re all watching the same video. That is what makes it a cultural phenomenon is that we all saw the same thing […] so letting people put up two or three different videos of the same thing at the same time kind of fundamentally breaks that, which doesn't feel right."

----------------

It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"

More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.

e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.

Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.

  • SchemaLoad

    We have just gone too far with engagement maximising. A/B testing titles and thumbnails was probably too far but we just keep going. At some point society needs to say stop, tech is addictive enough. We maximised way beyond what was reasonable and we need to wind it back.

    Give Google, Meta and the rest an award showing they beat human psychology and now we can encorage people to put the phone down and go outside again.

    • sicktriple

      Any lever by which it was once possible to impose "stop" is now completely bought and paid for many times over in the US. There is simply no way in which popular sentiment can meaningfully be imposed upon our society.

    • jocaal

      "We" have not gone anywhere. The rise of tiktok is a clear example that the network effects of the social giants are not strong enough to keep competition away. This is what the market demands.

      • SchemaLoad

        The market also demands meth and all kinds of things that are bad for people. It's time to seriously treat social media like the drug it is and ban all addiction mechanics.

        Algorithmic recommendation engines, infinite scrolling, notifications for suggested posts, daily rewards, all of it needs to be banned.

    • georgemcbay

      > We have just gone too far with engagement maximising.

      Data-driven optimization in general, and not just when it comes to things like "online content".

      Like, it certainly benefits us to a certain point, but after a while it starts to create fragile systems and we're getting deep into the fragile systems phase.

      See: the complete breakdown of the supply chains of just about everything, algorithmic price discovery contributing to out of control inflation, etc. This is beginning to impact nearly every aspect of our lives and IMO rarely in a good way.

    • hn_submit

      It will never stop as long as good money is being made with it. If you want it to stop simply stop watching.

  • CM30

    The point about political views made me think of a similar situation I hypothesised based on this announcement. Creators using A/B testing to decide which viewpoints most appealed their audience.

    Like I can see a video comparing two games getting variations where each option wins, then the creator using A/B test data to decide which of those views most appeals to their audience. Feels like a slightly sad way for creators to literally pander their views to their audience, regardless of their own beliefs on the matter in question.

  • doginasuit

    Can't be exposing people to ideas that under-radicalize them, they might go somewhere else. We are worried that AI is going to destroy society when ad companies are already well on their way. Attention at any cost.

OkayPhysicist

I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical. If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all. Users don't want their shit changing all the time.

  • sfRattan

    > A/B testing users without their knowledge and enthusiastic consent is unethical.

    Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.

    It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.

    So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."

    • sceptic123

      I think this is a bad analogy (or a good analogy for bad A/B testing).

      I would say it's more like: imagine if your dashboard controls and icons, had occasional tweaks in size and shape that made them slightly harder/easier to use, but over time ended up with controls you found more intuitive and easier to use.

  • cortesoft

    Hmmm, I am curious about which aspect of this you find unethical.

    Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.

    Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.

    Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?

    I guess I am just confused by this statement:

    > If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.

    This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.

    • xg15

      How about actually asking the users instead of experimenting on them?

      > or do we want companies that get feedback from users on whether new changes are helping or hurting.

      You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.

      • cortesoft

        Asking users for feedback is notoriously bad at generating good feedback. Most people don't respond, and those that do ask for things they don't actually want, or are only wanted by very few people.

        • xg15

          I've heard we have this amazing new technology now that can understand natural language and automatically convert it into structured data if we ask it to. Maybe we could use that to ask people at scale now.

          > and those that do ask for things they don't actually want, or are only wanted by very few people.

          Then how do you know what people actually want?

          • Joker_vD

            By observing their actual behaviour. It's notoriously different from what people claim their behaviour is.

            • isabelc

              Just because people do an action doesn't mean they wanted to do it.

              • cortesoft

                But you can’t build a business on “wants” that people don’t act on.

                If everybody says they want foo and don’t want bar, but when you make foo they don’t buy any but will buy a lot of bar, then are you “failing to provide what people want” if you just make bar?

      • fragmede

        > If I had asked people what they wanted, they would have said faster horses

        -(not) Henry Ford

    • carljungslabtek

      Obviously phased rollouts are fine and even necessary like you said. That’s not really “experimenting on people”, it’s experimenting on the network/system.

      I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.

      Profits are up though! In a big way! And our users continue to hate us more and more.

    • linster

      If I paid for a copy of Windows 95, I'm expecting Windows 95 in a box.

      When I have two free hours to turn on the XBOX for the first time in half a year, I want it to turn on right away and play my game. I don't want the box I paid for and have been looking at to figure that it needs an OS update, and a game update, and that the game should now be slower and glitchier than it was the last time I played it.

      When I play music on my phone in my car, I don't want to find out that the "Start Mix" button moved, or that showing the upcoming playlist now takes one more swipe, or that the UI won't load because YouTube Music doesn't cache it's UI anymore and when you have a cell network reporting 1 bar but it's actually zero bars, you get a spinner for ten minutes.

      The six CD changer in my dashboard has worked exactly the same way since 2006, the discs play when the key turns on, and nothing moves. There's no engagement to be had other than "my music plays when I turn on the car in my driveway which also has spotty cell service". There isn't a KPI to be measured, a PM to be promoted, or anything. It's just a car radio.

      Most consumer goods are solved problems. Nothing's changed since 2015. Even tech from 2015 is just a convergence of 2005 tech, like MP3 players, digital cameras, and Blackberries. People don't have radically different problems to solve in their day to day lives. People take pictures, share them, do email and group chat, voice calls, read the news, watch TV, pay for parking, do some banking. Watch a 90s TV show and all of those activites required different physical places and tactile goods. (Hell, that's why screens in Android are called *Activities*.).

      When I pay for a product, I expect to be paying for a finished product, not some psychological experiment that's someone else's promo packet.

      • cortesoft

        Most of the websites people are talking about here are free, they aren't things you paid for, so the argument that when you pay for a product you should expect something doesn't really apply.

        • spaqin

          Why not put that into question anyway? I paid with my data and browsing habits. Why are we wasting resources on solved problems anyway?

    • alpaca128

      > This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want

      Are you not aware how unpopular practically all recent changes on YT are among its users? Almost none of the changes done on YT in the last 5+ years would have happened if they took into account what users & creators want. So how exactly does telemetry and A/B testing help when it either tells them the opposite of reality or they simply interpret the data however they like anyway?

      • cortesoft

        I am not saying you are wrong, but how are you so sure that you know what the average user and creator wants? There are millions of youtube users and creators who aren't participating in whatever forum you are basing your information on. Maybe the people you hear complaining are a vocal minority.

        It could be that YT is just making everything worse for everyone, but I also know they have data that you and I don't have on how people actually use their product. I don't think we can assume they are just bad at making a product just because all the people we talk to agree with us that it is bad.

        • alpaca128

          YouTube's goals are not aligned with what is best for the majority of users. They are optimizing for maximum time spent watching ads, hence they are working against user interests.

    • tjpnz

      If I'm gathering participants for a similar study at a psychology lab I would be required to get their consent first.

  • BeetleB

    If you eliminate A/B testing, you'll get "A" testing ;-)

    > Users don't want their shit changing all the time.

    Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.

    You might as well just say "Updating an online service without asking the user first is unethical."

    Oh, and how much are you paying for that service...?

    • alpaca128

      > Oh, and how much are you paying for that service...?

      I paid YT Premium once. It unlocked a playback queue in the app. That queue had 5 separate bugs I found within an hour. It's literally a simple playlist and yet not a single feature (adding, removing, reordering etc) worked reliably. The playback queue also randomly emptied itself sometimes. Meanwhile I get a superior version of this in the browser by simply opening a video in another tab, for free.

      Why should I pay for that while they only have AI support designed to never solve any issues, do nothing against bots, and then warn you that you may get banned when you report too many bots?

      • jordwest

        God that playback queue has been so awfully buggy for so long. Rant incoming.

        I actually discussed it with a friend a few years back while debating the declining quality of Google's engineering.

        The craziest thing to me - it appears to be using some kind of eventual consistency, so additions/reorders/deletions have to go through some complex process server-side that takes several seconds to update in the UI (and often the order is wrong, or silently fails to add videos). And yet, the whole queue disappears without a trace if the YouTube app gets unloaded by iOS, and is unavailable on other devices, so it could have just been stored locally all along.

        I thought that perhaps they were storing it server-side because eventually it would allow the queue to be restored or transferred, but it has been about 3 years now and I don't think that day is ever coming. I just lost a whole queue of several vids this morning (on the plus side it was a good incentive to get off YT, so I'll give it that).

        My guess is it's using eventual consistency or some complex multi-microservice chain of RPCs because that's just what you have to do at Google. I'm sure there are engineers who want to fix the feature or go back and complete the rushed launch but likely can't convince the decision makers that it's necessary.

        So we get a subpar experience from one of the largest companies in the world with thousands of the best engineers, while Google keeps getting their $16/month because there are no alternatives

        • BeetleB

          > So we get a subpar experience from one of the largest companies in the world with thousands of the best engineers, while Google keeps getting their $16/month because there are no alternatives

          When it comes to paying for video streaming, they are plenty of other services, and they all provide a better user experience than Youtube. Even if I pay, say, Paramount for their ad-laden option (the cheaper one), it's superior to Youtube without ads.

          Why people pay for Youtube is beyond me.

    • dghlsakjg

      In the case of youtube? I am paying quite a bit actually.

      Updates are one thing, changes made specifically targeted towards manipulating users into spending more time on the site in ways they can’t opt out of (youtube shorts is a good example) are different. No one is bothered if YouTube updates to allow 8k streams. I am extremely bothered that YouTube does not allow me to disable shorts in the app. I don’t want shorts, they are a distraction and one more thing i have to guard against getting sucked into. Let me use the app how I want to use it, not how you want me to use it.

    • pixl97

      The above post is more evidence that software developers are unethical.

  • arcanemachiner

    I think an important distinction being missed by those relying to you is that you seem to be saying that A/B testing is unethical when it is used to optimize human resource extraction (i.e. marketing).

    (I have this weird feeling that you're not upset with blue/green deployments...)

  • akst

    > I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical

    Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.

    The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.

    > If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all

    This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).

    An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.

    A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.

    > Users don't want their shit changing all the time.

    Yep that’s why you don’t ask them.

    I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.

    When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.

    If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.

    • jordwest

      > This is a bit hyperbolic, empirics is something that should be used more by decision makers

      In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.

      > If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.

      There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.

      Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).

      Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.

      It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.

      A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.

      • akst

        Hey Jordan hope you're doing well.

        > In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough.

        With things like A/B testing, its not entirely an objective as you need to make assumptions which can be difficult to measure (although typically randomisation solves a lot of them), but you can only measure what you've decide to measure (which isn't random), so you don't know when you're in a local max. So IMO intuition and thoughtfulness is necessary. Sometimes product managers don't listen to data scientists when they say you can't measure Y with X, or the research design violates the required assumptions to make a causal claims (like reverse causality or controlling on a post treatment effect, e.g. employment as control when measuring income after hospitalisation (the treatment)). I think the worse offences I've seen have been from marketing teams.

        But proper research design does require intuition and thoughtfulness, because statistical models require thought, like other forms of supervised learning.

        I've seen both

        - PMs use questionable experiments to justify shipping something.

        - PMs dismiss experiments when it was a null result and shipped anyways.

        In either case I don't think the methodology is the cause of problems here, although I think shipping with a null result is justifiable if it's a larger unit of work (provided its not a regression).

        Sometimes things that have heterogenous effects get measured as a homogenous effect, Like say:

        - Your primary user base is X1 and X2 is a larger consumer base but makes up a small portion of your user base.

        - Your experiment does poorly with X1, but say there was an increase in user base X2.

        - However because X1 dominates the user base and your signups (because say you target ads to X1 over X2), no one drills into the effects on these different user bases, the result gets discarded as it seems to be a bad outcome.

        There's valuable information in the experiment outcome but without thought and attention you can miss it.

        > Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.

        I mean that sucks, but IMO with their market share, the way Google chrome is inclined to develop their product is very different to firms in more competitive spaces.

        Perhaps A/B Tests, allows google to optimise the things they are incentivised to pursue, but in the hands of smaller firms with different incentives are willing to tweak things to be more appealing to users when they have far less market power, which I think is probably more the issue in the case of Google.

        I just don't think this is a universal problem with the methodology

        • jordwest

          Oh hey didn't check your name before I replied! I assume this is the akst I know IRL, small world. Hope you are well too.

          > I just don't think this is a universal problem with the methodology

          In an isolated world I would agree with you, but in the messy reality we live in I think the broader problems with the methodology are two:

          1. The belief that everything can be measured. There are many intangibles (user trust, willingness to put up with bugs, "vibes") that aren't easily measurable. Yes, net promoter scores etc, but every company I can think of that uses them builds bland, buggy, largely disliked products that people use only because they have no choice.

          2. Focusing on experiments and measurable outcomes creates a tendency to de-prioritise anything that isn't easily measurable. For example larger, riskier projects that can't be quickly tested. Or whimsy - easter eggs that devs added to many products in the past that users remember for years. Rarely added anymore because they can't be justified against a roadmap full of experiments.

          Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside.

          To me maybe it's a bit of a question of science vs art. Take the video games industry for example, the AAA studios are largely moving toward the science end of the spectrum with predictable games that extract maximum engagement, while indie studios are mostly building small, unique experiences that don't try to dominate your attention.

          Experimentation always narrows focus down, and I think it's worth considering what's being traded off by doing so.

          • akst

            It's possible I've spent too much time looking at town planning where nothing is really ever measured and when numbers are produced it is it some insane

            Like claiming townhouses not facing out into the street somehow produces X $ in mental health costs due to "lack of inclusion" and the footnote links to a study where non-english speaking communities in Australia were having real health costs due to lack of access to translation services. Which is something that passes for "evidence based" planning. Doing experiments in social sciences is a lot harder tho.

            Maybe it's too easy to do experiments in Software and like you said

            > Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside

            I do think they're a useful tool but I see what you're saying.

            Part of me feel some of that is risk aversion, but also maybe some of its process dependence when you have a number that provides strong certainty on a number of things, and because there's that feedback loop they get drawn to the things that cause the number go up. Similar to how people become dependent on LLMs to get stuff done or affirm if they did the right thing, and when they're in a space that's harder to measure they don't know to judge if they've done a good job or not.

            I'm spending less time working on software, as I decided to get an econ degree, specifically econometrics, so I do spent a lot of time trying to think how to better measure stuff specifically in public policy, so I might be bias lol

            But I do get what you're saying, hope things are well for you

  • SaucyWrong

    Software changes before my eyes constantly via normal releases so I don’t really care if I’m part of an experiment or not.

    That’s software though—I see your point for something like content. I’m already used to seeing the title or thumbnail of a YouTube video change as the result of an experiment “winning” but the content itself…that would very jarring.

  • kypro

    Surely this is way too broad a statement?

    There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.

    Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.

    I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.

    • robobo96

      I agree. I work at a big insurance company with millions of online visitors each year. By A/B testing, we improve our services. A/B testing error messages has been a huge help for us. We can't ask users directly what they need: most users don't know what they need. That's why we do both qualitative (user research, online feedback forms) and quantitative (A/B testing, fake door testing, etc.) research.

csnover

From [0]:

> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.

[0] https://99percentinvisible.org/episode/673-the-score/transcr...

  • jimniels

    I read C. Thi Nguyen’s book “The Score” a few months ago and it’s phenomenal. Incredibly applicable to our modern age of metrics and measurements and organizations at scale.

    (And great quote, thanks for sharing!)

jordwest

As Seth Godin said:

> Enough A/B testing will turn any website into a porn site

  • dghlsakjg

    A corollary to the apocryphal Henry Ford quote: if I asked my customers what they wanted, they would have said “a faster horse”.

wewewedxfgdf

OT: YouTube has some big problems.

I don't think I am the audience for YouTube any more.

YouTube is converging towards where all the social media sites are:

* shorts

* photo/text posts

* longer videos too - typically 12 minutes approx

* AI videos - some of which are fine but I want to be able to filter them out

* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified

But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.

That "original YouTube" seems to be gone. Nothing has replaced it.

  • ivanjermakov

    Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.

    Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.

    • PaulDavisThe1st

      > Original youtube is there, it's called subscriptions tab. A curated list of high-quality content creators accumulated for over a decade.

      Alternatively, it's there in my RSS reader app, where I can see new stuff showing up without even visiting or interacting with YT.

  • denkmoon

    "The algorithm" is quite powerful. Most of my view time comes from long form content, between 30 minutes and 2hrs. Sometimes it can be a struggle to get it to show me a 10 minute video I can watch over breakfast or something. What Youtube pushes people towards is not uniform

  • hollow-moe

    They even added games lmao, all of them looks quickly sloped stuff. The only think keeping YouTube afloat is their inertia, everyone is here so everyone goes here. The day creators realize video is infintely copiable data and they can actually publish in multiple places simultaneously better platforms will emerge for sure.

ChrisMarshallNY

I'm not a fan of the new approach, which is, basically "We don't need to be creative, up front. We'll just shove some crap out quickly, and work on the friction points."

But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.

  • SchemaLoad

    A/B testing isn't about creativity, it's retention maxxing. It's what drives youtube face in every thumbnail, loud titles that don't say much, etc. This feature will optimise the actual content itself to be the most attention holding, overstimulating slop possible to synthisise.

    • jordwest

      > A/B testing isn't about creativity, it's retention maxxing

      Absolutely, and unfortunately in my experience many people implementing A/B testing actually believe that it is finding the "best outcome for users".

      I'm sure if we blind tested people to see whether they consumed more when unknowingly given cocaine vs protein powder, we would see cocaine win the A/B test.

smarf

> The skill of being a “creator” is not about maximizing views, retention, etc. It’s about finding creative ways to share stories, teach concepts, explore ideas, etc. Views, retention, etc. are all downstream of that.

this seems wildly naive

  • rglover

    That's exactly what the focus should be. It may be naive in a modern context where everyone is addicted to "number go up," but it's the only way to get an authentic, not manufactured culture.

    • ksmsisjxkwdj

      Hardly. If the top ranked ones were always the most creative, honest and there-for-the-art creators and artists, we wouldn’t have the Taylor Swifts of the world along with the tidal wave of fast-food content diarrhoea we currently have to suffer through just by opening a web browser.

      Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.

      Unfortunate. Depressing.

      • PaulDavisThe1st

        It might be a valid strategy for "growing", but "growing" is not necessarily the goal of all creators in any medium, certainly not in the sense that you mean it.

  • Rapzid

    Yeah.. Marques makes good content and puts a lot of effort into it. There are a number of Youtubers like that..

    But the vast majority are engaged in attention baiting; outrage, grievance, conspiracy, FOMO.. Take your pick.

    Even some oldies like GN are churning out grievance and drama.

  • MisterMunchkin

    It’s also ironic because he’s a slop producer

miladyincontrol

People who use DeArrow keep winning. For those outside the loop, it crowdsources alternative titles grounded on the video's actual contents rather than clickbait, and similar with thumbnails.

Presumably such tools will also force version A if this continues to be a thing. And likewise I imagine more minimal extensions will arrive to do just that one job.

jasonjmcghee

> YouTube spends way too much time chasing their competitors

Is A/B testing a thing on TikTok and/or Instagram (or similar platforms)?

sicktriple

Mainstream software development has been gutless and captured by MBA culture for decades at this point. Just like all other aspects of culture: video games, music, television. We can only optimize for one thing, money, and everything is downstream of that. It seems quaint to even point that out nowadays, since we've been living in the rubble of creativity for so long.

That's not to say these things don't exist but as always you need to go out of your way to find them, because people who are prioritizing the creative act typically aren't going to be prioritized in our attention-focused slop world.

skyberrys

What a strange choice from YouTube. They won't even let you create a draft of a video ( private not published anywhere ) and then replace the video before publishing. This seems much more confusing, it's like two videos yet they don't have different urls. Maybe they will control for a level of similarity?

jaggederest

Could go even further and produce multiple sections and let the "algorithm" edit them together in whatever order.

intro (2x 15 second clips pulled from random places in all the other sections)

section 1 (a/b/c)

section 2 (a/b/c/omit)

section 3 (short/long)

section 4 (a/b/c)

It would be like a choose your own adventure video, without the choosing, or the adventure.

  • PaulDavisThe1st

    This was already done for a documentary on Brian Eno that came out a couple of years ago. Nobody who watched the documentary could be sure that they had seen the same as anyone else unless they watched it together in the same physical location.

    • jaggederest

      It's funny because that strikes me as a perfectly fine thing to do for a Brian Eno docu, but for virtually every other program on the planet as absolutely egregious. I guess there's always an exception!

akst

In most cases the A and the B in an A/B test really should otherwise mostly identical, but one thing is different.

If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.

  • sceptic123

    Yes, lots of comments that seem to not really understand what A/B testing should actually be. I expect the implementation of A/B testing of video content on YT is unlikely to be done well either though so maybe it's correct criticism for incorrect reasons.

    • akst

      Yeah I don't expect most youtubers to understand that, and much of the guidelines may seem to insist upon themselves without understanding the underlying motivation for the whole approach.

      It does sound like youtube will be requiring experiments to pass some sort of similarity test before they can be run as an experiment.

      I think it'll be interesting to see how it goes.

trymas

I am probably looking waaaay too much into this, but I was told 5+ years ago that streaming platforms will develop AI to plugin "natural" product placement for each user.

For example, ad algorithm will draw product placement with specific products that it thinks you're the target. Protagonist will drink pepsi/coke/juice, will use linux/mac, will drive MB/BMW/Ford - depending on what ad algorithm thinks you'd like to buy.

Somehow I feel like this "A/B testing" feature is actually a cornerstone by YouTube for AI driven media future, for exactly this (dystopian) reason.

neals

Anybody else seeing ghosty black lines on hn after reading that dark page? Like a negative version of some sort, lingering in your vision?

fn-mote

Welcome to a world where commenting on a YouTube video is the same as writing a product review on Amazon. Now you need to include enough details that future readers will be able to tell that you are talking about a different product/video.

jamesforestwest

I’m not particularly biased against technology, but even to me, this sounds very dubious

j45

The comments here about A/B testing are one thing, the real wave that will be massive is unprecedented personalization compared to what exists today.

  • pixl97

    We all get our own little mini-socities with no one else in them.

    • QDwQ1

      may as well just cut to the chase and give everyone a fully automated heroin IV. that would maximize total happiness, after all

    • j45

      Very possible, I think in other areas it could be helpful.

      It's critical that specialists who know whatever they do, work with AI to see the path ahead for their area, while the mega models try to be a single mega model for everything.

tonymet

recently i was suffering in a "no reviews" experiment on Amazon, and it took a few moments to overcome "you're crazy it works for me" . Despite all of us building AB tests, few have embraced the ramifications that we are all experiencing a different combination of experiences.

Barrin92

Completely horrific especially in the creative context. If you make something, you are a creator, and what you make should be an artifact of you. The idea to A/B test a video, as if you don't know what the best video is you can make, to me completely undermines the point of making anything at all. At that point you may as well outsource your 'content' to a marketing department. You evidently only care about clicks, not expressing anything.

And even in the software world where there's a place to test say accessibility and whatnot this attitude is so prevalent that everything looks the same. Nobody's making a website like Larry Wall any more[1]. Can people please start making things out of their own volition again instead of following this brain dead attention economy

[1]https://www.wall.org/~larry/

smugengineer69

I've seen this argument a few times and honestly I don't get it - how do people propose we otherwise measure relative success of particular interventions -- Vibes? I've found that people who complain about how everything needs to be measured and how measurements are imperfect are often just bad at devising measurable quantities as proxies of real-world benefit. Sure it can take a bit of time to devise such measurements, and they're not inherently valuable, but if I can prove that a given variant performs better, why not pursue the preferable alternative?

It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.

It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.

If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.

  • liquicity

    So now the average Joe has to make 3 vids instead of 1 every time - as everyone else is now min maxing their engagement?

    Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection