Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

I suspect the tracking in VSCode is mainly used to improve the product. It's probably in my best interest, and the interest of the community as a whole to leave tracking on. I mean, I get it, HN is usually a more skeptical and security-focused crowd. At the same time, it's likely MS will just take that information and tailor their bugfixes and features to the things I need most, so by all means I want them to have it.


I think you’re right, but it needs to be a user choice. If presented with an option, I would likely leave it turned on 90% of the time.

I hope this project motivates MS to open source the runtime and making tracking a user- selected option.

I really like Code and think MS has really helped the dev community by making it so great and free. But I would like to see them embrace f/l/oss for all their non-core products and stay on track for customer/dev-friendliness.


Isn't it? When I opened VS Code on a new computer for the first time today, a message popped up with step by step instructions to opt out


Tracking should always be opt-in, not opt-out. I base this on expected user preference and trying to optimize for user happiness and functionality. I don’t think this is as true for non-OSS software where the purpose is likely more toward revenue generation than community and user functionality.


First, please let's draw a distinction between different kinds of tracking. We're smart enough to appreciate the naunces here. Some information is more valuable than others.

* How many times you clicked Edit -> Paste vs Ctrl/Cmd + V. This is not personal information. It cannot be used in any way to identify you.

* Your location/medical information. Guard this with your life if you must. Share it with literally no one. If it's uploaded somewhere, delete it asap.

All our information falls on a spectrum between these two extremes. Let's please acknowledge that not all information is super sensitive, that it's ok to share information like the paste example.

Second, this information can be useful, and can change the direction of the products we're building. For example, the Office team wanted to re-design the old menu bar interface to surface the features people were using the most. You know what people used the menu bar for the most? EDIT -> PASTE. You'd think everyone knows about Ctrl + V, but apparently not. The vast majority of the world used to click Edit, then click Paste. Knowing that users were doing this allowed the interface to be moulded to suit the needs of the silent, non shortcut using majority instead of HN power users - https://imgur.com/a/WLm4UJd.

Third, if you acknowledge that this information is useful, then it follows that it's only useful when it's opt-out. If it's opt-in and only 0.01% of your users opt-in, you can't make any reasonable conclusion from the data because it wouldn't be representative. If you tried to convince folks that no one uses keyboard shortcuts to paste based on this tiny data set, they'd laugh you out of the room. Collecting opt-in data in this case is almost useless.

Fourth, products that collect this info become better. If you're ok with webapps collecting this but not native apps (like VS Code), then the consequence might be that web apps end up being much better than native apps. Whether you prefer that or not is moot, but personally I like having options. Let's not force native app makers to stop collecting non-identifying information when the downside is minimal/non-existent.


With machine learning they might only need to know your pasting habits and what function names you use most commonly, maybe a bit of information about how fast your type and this will be enough information to identify you based on your behaviour.

Search engines are recording really detailed information as you type, as you browse other websites using their invasive trojan-like analytics. With this data you can probably identify anonymous users from their behaviour (even if the ipv4 address is "encrypted" hah). All it takes is for data from one company to end up at another company and I believe this happens very often.

I think the problem is that you need less and less information to personally identify someone as you have more data. We are not intelligent enough to know what data is identifying or not so therefore the only option we have is to stop giving out our data.


Or you could use differentially private data collection.


This is borderline paranoid IMHO


> With Machine Learning

Tell me, who exactly is running ML on this data? They're going to great lengths to anonymize this data. If they wanted identifying information, they'd just upload your unixname, they wouldn't anonymize and then de-anonymize with ML.

I don't even know how you conflated the metrics on menu clicks with uploading paste data. That's ridiculous. Did you even notice you made that leap?


If the menu click data is detailed enough then it may already be personally identifying information. You could take tracking data from another software that collects similar information and using ML match these together and suddently you can get a full name from just a bunch of seemingly innocent data (if the latter software is associated with a name of course).

I don't know what data they collect now or what they will collect later so I am just speculating. But I am sure the ToS has a section about how they can change the terms however they wish.

Please tell me about the great lengths they go to anonymize this data because I believe it to be very difficult to do, and absolutely not in their interest.

Edit: With pasting habits I meant if you paste with shortcuts or go through the menu, not the actual paste content. Sorry for the confusion.


> Fourth, products that collect this info become better. If you're ok with webapps collecting this but not native apps (like VS Code), then the consequence might be that web apps end up being much better than native apps.

So what? VScode can also be used as web app already, it's not about what others do or don't with their online services. The question is about the expectation and respect to the end user, having telemetry on by default is disingenuous to say the least.


How many times you ctrl + c & ctrl + v, and generally your computer usage, likely could be used to profile you.

One single link to your identity, and you now have user tracking. If course, there's probably easier ways to achieve it.


In theory, it kind of is: as far as I’m aware, which I haven’t delved deeply into it so grain of salt time, but it doesn’t send data until you make that decision. And it isn’t like you have to dig for it: it’s presented to you explicitly. Yeah, it asks if you wish to opt out of “making the product better” or whatever the verbiage, but I don’t think it is going to confuse many. Especially within the target user demographic.


What is the crucial difference between a popup saying "would you like to opt in?" and popup saying "would you like to opt out?"


Opt-in telemetry likely won't serve it's purpose.

Very few users change the defaults. To the point where you don't get enough data for useful insights.


That is exactly right. Almost no one would turn it on if it were off by defauly, which makes data collected by those few users strongly biased against general users and their usage patterns. (Meaning that the data you do collect will be useless.)

This is literally the only reason that telemetry is almost always on by default. There is no illuminati secretly buying developers to learn your secrets via application telemetry, and it's laughable to think that's the case, in my mind.


If people turn it off when they have a choice, and only leave it on when they are unaware it exists then obviously there is a problem. I don't see how it is laughable to be paranoid in this case.


isn't paranoia by definition irrational?


I suppose not when it is rational to be paranoid.


Rational people are cautious, not paranoid.


Opt out is when the tracking is on by default and the user has to take action to stop it. Opt in is off by default and action must be taken to enable it. What happens before the popup decides which one it is. If the user closes or ignores the message about tracking there should be none and there should have been none before the popup.


User perception of a default state.


There’s a good book called Nudge by Richard Thaler (and someone else I can’t remember) where he goes into great detail on the default decision.

Tl;dr; people usually accept the default so your prompts yield very different results.


You can't have your cake and eat it too. Telemetry in this case helps in improvement of the product, it's good for you as end user. If you don't like, please switch to another editor. You paid zero but setting requirements how it should and how it shouldn't. Good luck with utter ugly but no tracking and "free" GNU Emacs with its ancient Lisp and garbage packages.


You are the product. You paid by using it and by giving back on github or by getting your boss to buy support.

If you don't like that switch to a non-open source editor.


When I was using azure data studio the opt out was to visit a link, read the wiki page, open the settings.json file using your file manager and then paste in the tracking opt out key and restart the program. Or you could just close the message and be tracked. This is absolutely unacceptable opting out of tracking should be made exactly the same effort or less effort than opting in.


Wait a minute, there is tracking involved in a code editor? No, it's not in anyone's best interest. It can't be. Any info collected with tracking could also be used for other, malign purposes, covertly and illegally. I understand your perspective, I wish it was true as a fact. But it isn't. What I've learned in the recent years - don't trust anything that tracks and collects data unless it's your own thing.


There's a pretty wide area between a skeptic and a conspiracy theorist, but this kind of mentality certainly leans toward the latter. Logic dictates it's vastly more likely the data collected is to improve the editor, than a multi-national multi-billion dollar company collecting data to use it illegally.


No, it's an issue of trust. I wasn't implying that Microsoft does such things. Such things could happen without them even knowing about it. With security in mind I just can't take the risk myself or expose my clients to such risks. VSCode is obviously a consumer product, not an enterprise solution, so I advise my clients to use more trustworthy software, without built-in tracking capability.


Great-grandparent OP was most accurate in that this crowd cares most about privacy regardless of what it is used for, and that it would have little utility outside of this crowd, but your dismissive response is as if you've been in a coma for 10 years.

As grandparent OP said, in recent years this has moved waaaaay beyond theory territory and been shown time and time again that <Corporate Sector> + NSA + FBI + CIA + intelligence agencies around the world all employ different ways of collecting analytical data broadcast over the internet.

NSA just taps the servers without telling anyone.

FBI sends National Security Letters containing gag orders preventing companies from telling you that the federal government is now a data sharing partner.

CIA just pays companies for it.

FISA court issues secret rulings justifying the legality of it all.

Whether that bothers you or not is up to you. Most people don't care. I usually don't. I wouldn't say "logic dictates this won't happen" when it is pretty much only the multi-national multi-billion dollar companies subject to this kind of tampering, and most incentivized to monetize the analytics by allowing these and unknown third parties in.

To avoid the security exhaustion, some people would simply prefer their text editor not be "smart", which is a euphemism for internet connected.


Google collects far more data than Microsoft does. They have zero problem processing it all.

It’s not like it’s human sifting through it by hand.


> Wait a minute, there is tracking involved in a code editor?

I recall a recent thread here on HN where I learned that the Windows Calculator was sending data ("telemetry") back to the mothership.

Nowadays I think I'd be more surprised if a Microsoft product were released that didn't phone home.


One of the more flabbergasting comments I've seen is when Mozilla removed ALSA support from Firefox because nobody stepped up to maintain it, and the metrics showed that it wasn't widely used. There were people complaining about it being removed, and stating that the people who did use it were more likely to have turned off tracking, and thus did not show up.

I mean, sure, it might be the case that there were indeed tens of thousands of ALSA users with tracking turned off, but... From my perspective, it seems more likely that it was just a handful, and really there's no way to tell the difference. If you turn off telemetry, and be aware of and accept the downsides.

I trust Mozilla, so my telemetry is on, but for many other applications, I often opt to turn off tracking - with the understanding that it's harder for them to tell what my needs are.


So now we need to allow surveillance, just to prevent the biased data set from ruining the software even more?


No, you don't need to do that, but if you don't, you should realise that that results in biased data, and thus influences the focus of development. If you have an alternative other than "the developers should just magically guess what would make users happy", then I'd be happy to hear it, but otherwise, that's just the way the world works, and thus that's the trade-off you will have to make.


Then less people will use there new version. If they want to rely on this one datasource only be prepared to have it gamed or not reflect reality.


Well, again, it'd be great to hear what other information they should rely on other than magically guessing what people want.


It maybe starts with simple statistics. But then you want to know what features the user use, then you want to know what other programs they have installed. Then you want to know what the users search for on the web. etc. It's a slippery slope.


Why would Microsoft gathering usage data about VSCode turn into spying on the user's other apps and web searches? There's no connection between the two. Tracking usage of VSCode features has a clear connection towards improving VSCode. Spying on the user's activity outside of VSCode has no connection at all to improving VSCode.


Every time I hear the phrase "it's a slippery slope" uttered by someone arguing against something, I am immediately suspicious of the argument that person is making.

There really isn't such a thing, in the way you've used that phrase.

Capturing telemetry on how I use a tool from within that tool is perfectly fine, to me. Collecting telemetry on my search history in the browser by that same tool isn't. THERE ARE NO INTERIM STEPS that makes the second of those ok. There is no slope. If there is, it isn't slippery. There is a series of discreet decisions and at some point (which is different for everyone) a line is crossed. There was no slope or slip that brought you there, only a series of mostly unrelated decisions.

To think that Microsoft's long-term goal is to install a keystroke logger via a multi-decade and multi-phase plan that begins with application usage telemetry in a free developer tool thanks to "a slippery slope" is just simply not realistic.


When you walk in the wrong direction the final step off the cliff is the last one. Better to get off of the slope because choices get fuzzier the closer you get to the sun.


It's called a fallacy for good reason.


Im not making that up, it was in the TOS for VS/code last time I checked.


You've literally described a slippery slope fallacy


It's not a fallacy. When a printer driver reporting stats back to the manufacturer made national news in the early 2000s and prompted calls for laws regulating privacy, the argument was that it was a slippery slope fallacy to assert that privacy violations would get worse. Look at where we are now.


>>It's not a fallacy.

By itself no, but it is often used fallaciously. Such as in this case, when someone is opposing some good thing on the basis that that good thing might, some day, lead to bad things.


Not realizing that many things actually ARE a slippery slope is what has gotten the world into a lot of messes. Nearly every legitimate privacy concern that we have today started out with a legitimate and we'll meaning purpose.

Our entire legal system is predicated on common law precedent. So it is very valid in many cases to argue that allowing something good now, might set us up for something very bad later.


What you’ve described there is exactly how bad things come to be.

You think people elect for bad things? Bad governments, bad software or privacy violations? They chose things that are full of rhetoric and promises of good things then those bad things get snuck in off the back off relaxed regulation, or existing software adoption, etc.

Take Facebook as an example, people didn’t sign up to it thinking “I wanted to be tracked around the internet so I can have personalised adverts” nor dis Zuckerburg think “Wouldn’t it be good to create a platform that could latter be used for rigging elections”. No, instead we got there because of a serious of good ideas that slowly got abused.

There is a saying that goes “The path to hell is paved by good intentions.” I’m not a religious man but I think that beautifully illustrates how slippery slopes are not a logical fallacy.


Good things such as trackers ostensibly designed to deliver you relevant content?


No, good things such as telemetry and automated error reporting so that bugs can be fixed effectively and efficiently and everyone is better off for it.


That sounds nice. But could they use this data legally in another way?


"Slippery slope" is not a fallacy. Life is full of slippery slopes in terms of behavior.


https://en.wikipedia.org/wiki/Slippery_slope#Non-fallacious_...

No, that still falls into fallacious usage.

The user doesn't justify why the steps of their assertion follow after the other. Just that they... do.


No one needs to justify the "why" when we're talking about a very well trodden slippery slope.


Fallacies don't just stop applying when it's convenient to elide justification.


That's true. However, we also don't collect history for the fun of it.

Over the course of the last 20 years, we've seen that once a data collection and digital surveillance framework is put in place, the surveillance tends to expand.

Slippery slope arguments, sans good reasoning, tend to be fallacies. However, don't fall into the trap of thinking that an argument backed by historical record is a slippery slope just because it's predicting an outcome. We might call that the "history is all slippery slopes" fallacy. Stating "this has happened before multiple times before, and each time has lead to x" is a very different argument to stating "this has happened, so the logical extrapolation is x".


if it can be that well trodden, is it really still a slippery slope?


Overwhelming empirical data across history.


That’s proof of existence of slippery slopes, not of prevelance of hypothetical slippery slopes being realized.


We’re talking about data collection and usage, not possibilities in the universe.


Exactly, which is why I'd be more interested in concrete commonalities of slippery slopes that were actually realized, and not just a general "history has plenty of slippery slopes".


The fallacy is to assume, without evidence, that the slope is slippery. There are plenty of slopes that aren't. Probably most, just you don't think about those because you know they aren't slippery already.

For example, I slept in 'til 8:30 today. OMG, a slippery slope. Next thing you know, I'll be sleeping until 3PM. Til midnight! I will never wake up again. But as it happens, sleeping in isn't a slippery slope. I don't think there's any solid evidence that telemetry is either.


Umm, it's slippery for me. I haven't had a stable sleeping schedule for over ten years. Send help.


This is exactly why I get up a minute earlier each day. As it is, I am waking up tomorrow at 4:39 AM. Mission accomplished.


(10 years into one possible future)

“... and that was me before we had children.”


That's the "hasty counterexample" fallacy.


I think it sends MS every keypress in search fields - https://github.com/Microsoft/vscode/issues/49161


If you had actually read that issue,

> Lol, probably just an oversight because they made the search a lot better.

> So I think it only does this on the settings file and not on other files ;)

> https://code.visualstudio.com/blogs/2018/04/25/bing-settings...


No matter what, it's an opt-out setting that can only be disabled with:

"workbench.settings.enableNaturalLanguageSearch": false

That's unacceptable.


Can you explain why this search feature is a problem because I don't understand?


Maybe not understanding the problem with a tool that may be used to work on proprietary code containing trade secret information silently and unexpectedly sending information out to the Internet is the reason for software becoming spyware...


Emphasizing from my original post,

>> it only does this on the (VSCode) settings file and not on other files ;)


Starting with "Lol, probably just an oversight" and ending with an emoticon is excessively dismissive and flippant... that's how you lose customers, or potential ones.


You store trade secrets in your configuration file? I still don't understand what the big issue is. Can you clarify that instead of critizising my lack of knowledge?


It's a shame we have to assume the worst. What would be better is for Microsoft to be more transparent about the telemetry and to enable more granular control. That said, I wouldn't bother checking as I don't really care. I rub shoulders with infosec issues daily, as most IT folk do these days. When I think about the perceived risk of telemetry from VSCode, now and in the future it's a negligible risk that I accept.


They are actually very transparent about the telemetry they collect, and they offer granular control. There is a log of all events sent to MS, and a page in settings dedicated to the different types of telemetry they can be enabled.


Where is the log stored? All I can find online are some instructions that seem out of date as they don't refer to the interface I see (or I don't understand them): "You can inspect telemetry events in the Output panel by setting the log level to Trace using Developer: Set Log Level from the Command Palette." [1]

The "page" in settings consists of just two options, and the complete descriptions of the types of information they collect are "crash reports" and "usage data and errors". That seems the opposite of transparent and granular. Am I missing something?

[1] https://code.visualstudio.com/Docs/supporting/FAQ


Windows has an overarching telemetry collection system and VSCode might use that on Windows.

There is an application in the Microsoft store you can install that lets you view all telemetry collected by Microsoft, and the actual data is encrypted. The metadata (which tells you WHAT is being collected) is not. Knowing the people I have worked with in the past, this is to securely prevent modification by users before the data is actually sent. I've worked with lots of people who would modify that data to attempt to get a feature added that they wanted or just to screw with MS.

That tool also gives you the option to delete all telemetry sent from that machine in Microsoft's possession.


Thanks. The app is called Diagnostic Data Viewer.


The command pallette is what you see when you press CTRL+SHIFT+P, and typing 'set log level' in that dialog will bring up a search result you can click to set the log level' (the level of stuff that gets shown in the Output pane).

Setting that to "Trace" will show the telemetry being sent in the Output pane of the interface, amongst a bunch of other stuff, I am sure.


FWIW, I'm completely comfortable with telemetry, analytics, whatever so long as 1) there's a complete local log and 2) I can opt-out.


Incorrect. That page has essentially no detail.


IMO they could get whatever statistics they wanted, as long as they asked before collecting


and also, what they intend to use said statistics for.


One could be totally uncritical about the current or future intentions of Microsoft and still prefer a version with no telemetry. Data breaches are very common. Employees, hackers, or governments could all gain access to the data.

The data could be accidentally broadcast or left in a vulnerable place. Even the payroll data for the national security establishment was once reported to be compromised. Everything is vulnerable. Computer science is in such an abysmal state.

Even with properly configured servers, OSes, and databases that are up to date, they are still vulnerable to zero-day attacks because they are not formally verified and have enormous and largely unnecessary complexity. Then throw in the crazy complexity of processor instruction sets, creative side-channel attacks, and stuff which exploits the physical properties of the hardware (rowhammer).

It is reasons like these why we should never really trust transmission of sensitive data over the internet. The concept of secure voting systems, for instance, is literally a joke. insert obligatory xkcd here


Yeah, there is a reason why the Kremlin uses mechanical typewriters.


you should read up on slipper slope fallacy

https://en.wikipedia.org/wiki/Slippery_slope


did you read the article you linked?

"Logic and critical thinking textbooks typically discuss slippery slope arguments as a form of fallacy but usually acknowledge that "slippery slope arguments can be good ones if the slope is real—that is, if there is good evidence that the consequences of the initial action are highly likely to occur. The strength of the argument depends on two factors. The first is the strength of each link in the causal chain; the argument cannot be stronger than its weakest link. The second is the number of links; the more links there are, the more likely it is that other factors could alter the consequences.""

https://en.wikipedia.org/wiki/Slippery_slope#Non-fallacious_...


Where's the link between "collecting data on VSCode feature usage" and "gathering a list of all other apps the user has installed on their system and all web searches the user does"?


I mean, I can definitely see “all other apps” being collected as a way to check if there are apps conflicting with or interfering with VSCode. Maybe it’s only collected for a small subset of users with particular issues - but it would still be tempting for a dev to try and collect.

I can also see them collecting code searches done within the app as a way to check if their search system is working well for real use-cases.

Neither is outside the realm of possibility - you just have to put yourself in the mindset of a dev who is assigned to track down a rare crash or to “improve the search experience” who might want a little more data to work with.

Not saying I agree with any of this collection - it’s terrible and definitely falls under “the road to hell is paved with good intentions”. Companies should be extremely clear about what they will and won’t collect - and never cross the line even if it would be useful.


Even with the best intentions at heart, do you trust them not to accidentally leak sensitive information about you through their telemetry? There are so many places a "phone home" system could either be compromised, or accidentally send more than it should.

Telemetry is used either with naivety or malice. There is always some risk to the user.


The telemetry code is in the source. You can look at it yourself. It's anonymized.


I completely agree and I’m in the same boat (I’ll continue to use VS Code).

That being said, it’s somewhat amazing we are now in a time when a Microsoft product could have aspects users don’t approve of and rebuild it without it: everyone wins.


Are the telemetry data made public? If it's not, then it could be manipulated to justify arbitrary decisions from MS


Why would they make arbitrary decisions that run contrary to what the telemetry data shows? That just doesn't make sense.


To gain cheap trust from people by saying "trust us, the data shows it"

I can think of two reasons for the data to be made public :

- for trust and transparency purposes : It is for the same reason than an election system should be observable and reproducible, from data collection to final decision, including counting methods, etc. Otherwise it would be like "trust us, the data shows it" and you don't show the data to anyone to prove it.

- for coordination and sharing knowledge : some people might interpret the data differently, chose to focus on a niche market by doing different bets than MS, and create a complementary editor to the one from MS. MS has no obligation to support minorities, but someone else might be interested, and those minorities are detectable in the data


This may be how they develop Skype hah.


Yeah, if you're going to just do what you want, you don't need to collect telemetry to do it.


I'm ok with until its get used to improve the editor. It would much better if they can share the data they are collecting.


You can view a real time log of all telemetry events in-editor. Or do you mean you’d like them to share with you all he telemetry from all their users?


Well, why not? A bunch of people here are more than happy to defend Microsofts tracking. If telemetry really isn't a big deal, make all that data public. It's our data anyway, collected from how we use their software.


That would be like publishing the recipe to the secret sauce that gives them their competitive advantage. I doubt that theynwant to share that insight.


It’s possible that MS only uses it for bug fixes. But it’s also known that MS shares data with the NSA and other government agencies. It’d be very unlikely that the NSA would say, “yes, we want your data, but not Visual Studio data. That’s private. :)”


Does anybody else get the sarcasm?


> tracking in VSCode is mainly used to improve the product.

Not only, that's the main problem. And I can't simply trust it for company without dark M$ reputation and EEE experience.


> It's probably in my best interest, and the interest of the community as a whole to leave tracking on.

I hope this is sarcasm. If not, what did I miss? This is the same empty phrase that facebook, google, etc. use. Why is Microsoft more trustworthy in that regard? I am 100% sure they use the data to make money in short terms or in a long run. They for sure use it to make VSCode better, but only to get more people use VSCode and make them dependent on it. VSCode is a prime example of the Embrace, Extend and Extinguish strategy. I already see them grasping for the Python community.


I agree. As someone who builds product (with opt in tracking) understanding users is key to building a good product.


Absolutely, also it is amazing how little tracking does to understand users.


So then why keep the telemetry source code secret?


The telemetry source code is not secret. In fact, it is annotated thought the code, in order to comply with certain GDPR requirements. Try searching "GDPR" globally in the vscode source.


The point behind open source software is that the users themselves can make pull requests when there’s something they don’t like.

A need for telemetry indicates that contributing is too difficult.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: