RFC: Proposal for estimating unique Julia clients without tracking

Thank you for this official moderator clarification. Having the project’s governance standards explicitly placed on the record is incredibly helpful.

In your assessment of whether this environment is “toxic,” I notice you selectively focused only on the terms “insane” and “antisocial.” You conveniently omitted the parts of the comment that equated standard FOSS privacy principles with the ideals responsible for “shadow banking,” or the labeling of dissenting users as “foil hat extremists.” If project leadership considers that to be “nuanced pushback” rather than bad-faith hyperbole, the bar for civil discourse is functionally nonexistent.

Furthermore, pointing to the quote, “Julia is better when there are all kinds of people, including people with stupid political and social ideas,” and framing it as a genuine “invitation to remain” is a remarkable misread of a very basic rhetorical device. Tacking a hollow welcome onto a litany of personal insults is a standard tactic used to deliver hostility while maintaining plausible deniability. For a community moderator to endorse that tactic as proof of a welcoming environment is deeply revealing.

I do not need to argue this any further. By officially ruling that this behavior is acceptable, and even endorsing it, the moderation team has perfectly validated the core premise of my departure.

Thank you for making the reality of this community’s governance undeniable. I yield the remainder of my time.

Since you’re replying to my deleted post (which I might as well undelete, since the damage is done), I must clarify that I am not a moderator. I have absolutely no influence on the moderation or governance of this site. As far as I can tell, no moderator has officially commented on whether any behavior in this thread is acceptable – although you might, and obviously do, take the lack of a moderator response as a response in itself.

If I was a moderator (which I really wouldn’t want to be), I don’t know if or how I would have asked @jakobnissen to more finely thread the needle between calling out a set of bad and harmful ideas, and giving personal insult to people that subscribe to some subset of those ideas. But I do agree that we should try to avoid to cross the line to “character impugnment”. Whatever ideas you personally hold probably do not make you a bad person, or someone who shouldn’t have a place in this community.

Good! I’m very happy to hear that. I was pleased to realize that it was a very simple change to the protocol.

We keep raw logs with IP addresses and everything else for 31 days for incident response, but it goes into a completely unrelated S3 bucket that the longer term log aggregation system never touches. So we can start bucketing right away.

No header sent is equivalent to “opted out” so you can just do nothing.

That would be appreciated, just so we can know that from the logs. We could potentially add a header for it if that’s helpful.

This is fine, of course, but it’s also designed so that the client portion is very simple to implement. And I would be happy to help. Pretty sure an LLM can do it pretty well too, especially given an existing implementation and thorough spec to base it on.

That would be very cool and I’d love to see this help open source projects everywhere gain better insight into their usage statistics while preserving people’s privacy.

It’s unfortunate that you keep misrepresenting what’s happening to make it sound sinister. “JuliaHub is doing this” — when, in fact, as has been stated many times, JuliaHub has nothing to do with this and doesn’t benefit in any way, nor receive special access to any data. “Moderators have officially endorsed this position” — when, in fact, the statement wasn’t made by a moderator. Some of the statements in this thread have been spicier than they ought, but I think that observation goes both ways.

If you have objections, ok, but please make them based on reality rather than based on fabrications. It undermines the position you’re putting forward to keep insisting on arguing from false premises.

Again, I think this misrepresents reality — it asserts that something bad is being proposed that doesn’t match what’s being proposed, and then argues from that misrepresentation. I asked the following question on the pull request a month ago in response to a similar post and got no response:

So I’ll give it another try here: what, concretely, is the way in which this makes anyone’s local machine less private? Violating individual privacy is precisely what the protocol bends over backwards to avoid.

Is there some way that you believe the protocol fails to preserve individual privacy? If that’s the case, I’d very much like to know what it is. I think I’ve demonstrated that I very much welcome that sort of objection and am extremely willing to address any such issue. That’s precisely what the back and forth with @foobar_lv2 has been about, and it has led to several concrete improvements.

Is the objection that you don’t trust the review process and believe that there might be a flaw in the protocol that simply hasn’t been discovered yet? I have proposed that we can implement the client side but wait to implement the server side until the protocol passes peer review. Of course, that doesn’t guarantee lack of flaws either. However, if a flaw is discovered at some point in the future, we would immediately turn off the server side and remove this from future clients.

If the issue isn’t actually about these things then it doesn’t seem to be about individual privacy at all. If not, then what is the issue? Do you object to the Julia community having aggregate user data at all? That’s not really about individual privacy then, but it could be an objection. I’m really just trying to understand waht the actual issue is here.

I skimmed through the Executive Summary. I don’t like it, but I don’t like anything that collects data about me in any form, and I am generally opposed to any forms of gamification particularly in my private life.

However, if you and @foobar_lv2 are saying that it is bulletproof, and will remain so for years, if not decades, and that it will help individual researchers make a living, then I’m perfectly fine with it. Especially that the option for the opt-out is included.

On a separate note, to be honest, I am skeptical that this will significantly advance the development of the language or achieve the promoted goals. In my opinion, this approach is unlikely to disrupt established Silicon Valley technologies such as Python and C++. However, I might be very wrong.

Proposal for moderation @StefanKarpinski (since it’s “your” thread) and @mbauman (as the resident moderator) :

Let’s split this topic. One thread for technical considerations, that ideally leads to an implementation of HLL-RSA user counting to be merged as a PR, initially disabled by default (this makes the actual data pretty worthless).

A second separate PR and thread / discussion / community consultation for whether HLL-RSA counting should be enabled per default, i.e. switching from opt-in to opt-out, that tries to build rough consensus for that.

It is anyways good development hygiene to have separate commits / PRs for “support feature X, gated behind feature-flag” and “enable feature X by default”.

Most political considerations and forum flamewars go to the second thread. First thread gets ruthlessly moderated / redirected to the second thread.

You’re thinking too small.

If I want reader statistics for my blog, then your proposal is a good way to do this. There is a not insignificant population who (A) wants statistics on resource-using clients, aka readers, and (B) hates cookie-banners, and (C) cares about their user’s privacy, and (D) also cares about the general politics of user privacy.

That would require a small javascript implementation of the client-side protocol, putting the key-material (i.e. the user’s statistical-tracking-identity) into localStorage.

In a perfect world, the necessary machinery would be part of web-browsers, such that users don’t need to trust the javascript library (this won’t ever happen).

Wow, I didn’t think of this. That could actually be quite impactful.

While there have indeed been three or four concrete discussion threads (technicalities, opt-out-vs-opt-in, philosophies of privacy in general, and motivations here in specific), I don’t think a topic split makes sense nor do I think any are really off-topic. In fact, the motivation of the split, to have a place for a “forum flamewar” is misplaced: we do strive to avoid flamewars here. If anything the technical discussion makes most sense on GitHub itself.

The big challenge here, though, is that this topic explicitly wants for larger community feedback on the topic, and that feedback has included ideological concerns. Our community is large enough to have widely diverging views on the ideology, including quite strongly-held contradictory ones, and that’s ok! That’s precisely why we strive to avoid ideological flamewars; they become exclusionary fast and nobody is changing their mind on the internet here anyhow.