Hacker Newsnew | past | comments | ask | show | jobs | submit | mayakacz's commentslogin

I'm not one to comment often but this really pisses me off.

OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!).

Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to have some punk from OpenAI lie to you, threaten you, and tell you that they are willing to go on the record that you "deserved" it? What is this, the Godfather?

If OpenAI solved Navier-Stokes, that is an astounding result! - yet they'll still be remembered as those who thought credit was more important than results. That winning was more important than collaboration. If this is true, they're burning any trust left with academia.


I'm stunned that people are taking this accusation as a fact.

OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.

There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.


> The entire training data thing is speculation.

I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that.

I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-opt in without noticing.

All my work and conversations since I don't know are now part of their training corpus. No way to take it back.

Google's Gemini/Antigravity didn't have opt-out toggles at all last time I checked.

Codex also has a separate "include environments" setting which is hard to find (found it in Codex Cloud) and I don't know what it does.

Lots of Dark UI Patterns here even if we assume they keep their promise.

For this incident, Occam's Razor says their internal models somehow saw a version of the mathematicians' logs, during or after training. Maybe indirectly.

These systems are literally designed to collect data. Privacy and safety is not trivial to achieve on the users' side. Simply because it's against the labs' best interest.


You can opt-out of training in Gemini on personal plans, but it disables your chat history, just to be vindictive; there is no technical reason and the other companies don't do this.

But the whoreshippers of The Holy Dollar will tell you it's all good and justified.

He didn't even make that accusation!

> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.

Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.


Parse that statement more carefully.

> I was told the model did not look up user data.

The naive way to read this is "Nothing you guys did influenced the way our model got to the solution".

The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem".


Duh. There are supposed to be limits to what OpenAI is allowed to access with respect to logs and user interactions but there is no technical limitation.

It's a bit like sending unencrypted messages through a messaging app and the developer having a TOS that says they don't look at your messages. They might not, but they are fully capable of doing so. If they have a reason to do it, they will. Nobody's stopping them.


I read that as “the model didn't look up user data” as part of a “tool call,” i.e. they don't have an internal tool that loads user data (chats, sessions, attachments) for their internal models to read online while working.

Or (likely) they do have it, but the model didn't use it (unless it's so powerful it escaped that guardrail, wouldn't that be ironic?)

They declined to answer about anonymized aggregated user data being used for training. And even then, they may weasel out that they don't train on your “input” words, but that it's fair game go train on their “output” to your words.


>> Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.

How "fast" does it have to be? Buckmaster and Alpoge have been working on this for just a day short of a year. See Alpoge's tweet announcing his collaboration with Bukmaster dated 9/19/25:

https://x.com/__alpoge__/status/2097206973418611054

It takes a few months to train a model these days but not a whole year. OpenAI had all the time to train on Buckmaster and Alpoge's results of just a few months earlier at which point they must have been well on the path to their result.


It would be shocking if it wasn’t trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? It’s so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?"

There’s potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber?

The only unlikely part is the timeline - your sessions from a week ago probably haven’t made their way into the model. It’ll just take a while longer, and will be massaged just enough so that it isn’t really your exact session word for word so you can’t sure as easily.


at first I thought your post was a bit revolting with "have you read ToS?" bit, but in the end I completely agree and understand

I also don't get why it was downvoted, other than due to people not reading past the first sentence - although in the modern world's attention deficit that is understandable too


openAI's claimed solution uses a model trained in the last 2 weeks. The prior work would definitely be included in the training set.

And the labs are all building panel of domin expert models, while simultaneously chasing open math problems.

I would be shocked if they weren't tuning those models with the most relevant math texts and user material


It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course.

Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you?

The whole reason they have subs is to train them on YOUR WORKFLOWS lol


They are not training a whole model in a matter of days

They where working on the problem for a year using codex.

Models are very obviously continuously updated.

Model editing to remove PII that slipped through, all sorts of things of that sort.


pretraining is months but they can totally fine tune in a few days

They don't need to train a whole model. They can feed it new information and fine tune it.

Why would this be implausible?

ChatGPT user sessions were found publicly exposed to the internet not too long ago. Moreover, OpenAI has continued to play a hype-marketing game by revealing how their models keep breaking out of the sandbox.

Conspiracy minded thinking is not helpful, but why should OpenAI be granted the benefit of the doubt here after being caught doing underhanded/negligent shit on several previous occasion?


>ChatGPT user sessions were found publicly exposed to the internet not too long ago

Do you mean publicly shared chats were able to be accessed by the public? That's the point of the feature.


Couldn't the Enterprise have a different fine print?

Given the history of OpenAI and current litigations, I would say they've developed a bit of a reputation for not respecting intellectual property. I'm dubious they have some unbreakable moral code that would prevent them from viewing and using user data.

Everybody knows it's not a sure thing, it's a question of trustworthiness. OpenAI is not trustworthy at all; this random researcher is and seems honest so far. iThe fact that people are corroborating Bubeck being a piece of shit in other settings add to credence. But nobody is over here saying it's an indisputable certainty.

And your (2) is probably false, their history of deception suggests they would do just about anything as long as they didn't think it would backfire on them publicly.


This is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so it’ll now be treated as a fact and repeated endlessly in the echo chamber.

I am stunned anyone is giving OpenAI the benefit of the doubt

I'd be surprised if all of it is organic discussion, shall we say. I reckon The Bot Factory just possibly might dogfood the astroturf machine.

I guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max

He asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?

Whenever I see comments defending AI companies, I look at the account's creation date, and interestingly almost all of them were created post 2024.

Absence of evidence is not evidence of absence. With the behaviors we know OpenAI engages in the accusations are wholly believable.

> I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.

It's a gamble that this would be overlooked compared to the reputation they build for solving the thing


> playing around with user data like that would destroy their business.

Their entire business is based on stealing data. They can make a calculation that the cost stealing data is less than the cost of the positive publicity they can shape for solving Millennium NS


>The entire training data thing is speculation

Quite literally in the terms of use.


I don’t think you understand how brazen big tech companies are in practice.

They stole it.

Especially after the blatant cover up of their uncontrolled bot swarm infesting the internet, and the feckless "hopefully we do better" response upon being caught, I don't think OpenAI deserves much grace until they properly explain themselves.

We had all assumed that surely the supposed smartest engineers in the world, with access to the most computing and a direct view of model capabilities, would take sandboxing and cybersecurity much more seriously than they have turned out to do. It follows that while we might assume they take user data privacy seriously and have tight controls on who can access it, it's possible they do not actually do that.

At this point any initial trust is dead and has to be re-earned.


> I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.

And that's exactly why you can only see evidence of this when the stakes are high enough, such as solving a Millenium problem. The legalese they outputted in response to the incident left them an escape hatch that permits the possibility of theft. People familiar with corporate damage control should recognize the verbal maneuvering, often used to paper over actual guilt.

There was also already a high prior of shadiness. The company is run by someone who is close enough in reputation to "known sociopath".

It's actually far more reasonable to assume that OpenAI stole user data to generate a breakthrough. They have an unbelievably high economic motive to do so. And even a low chance that they might be doing this implies catastrophic risk to anyone with valuable knowledge.


This Godfather-like threat in particular pissed me off as well:

> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”


This is what is truly problematic. People at OAI probably think they can do and say whatever they want.

This conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data.

And simply knowing a problem can be solved is half the battle.


From Buckmaster's text:

    The route to the Clay problem through a
    smooth force, options c and d in Fefferman’s statement of the problem, is the
    route Luis and Diego opened and the one Levent and I had quietly chosen to
    attack. Almost nobody else I know of was working on it. It is not the direction
    one arrives at in a few days by giving a model the problem statement. When I
    heard “forced,” it was a bright red flag.
This is much more than the knowledge than the problem can be solved, it's also the specific, non-obvious approach to solving it. That's much more damning for OpenAI, if confirmed.

That's a stretch. The Luis and Diego paper was published in 2023 and is included in every frontier model's training dataset. An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and Alpöge’s work.

And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.

There is not enough information at this time to reach a conclusion. The best option is to wait for statements from both sides, then reevaluate.


> An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and Alpöge’s work.

the post you were replying to quotes Buckmaster specifically denying this: "It is not the direction one arrives at in a few days by giving a model the problem statement."

> And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.

"implies" is a surprising choice of word here. that's certainly one interpretation of "an insane amount of compute had been used". what came to my mind, considering Buckmaster's statement that the AI would not head down this specific path on its own, is, though, that they prompted it in this specific direction and then used an insane amount of compute. this seems consistent as well with these other statements:

> Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler (...)


>"It is not the direction one arrives at in a few days by giving a model the problem statement."

To be honest, he was referring to routes (c) and (d) to the millenium problem, as far as I understand no more specific. Which is 2/4 routes.


i don't know if i understand what you're saying. but i'm no mathematician. here's the full statement in question once again:

> The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.

so, you're saying that "the direction one arrives at in a few days" is (A) "the route through a smooth force, options c and d in Feffermann's statement of the problem", and no more specific than that, i.e. does not necessarily include (B) "the same path route as Luis and Diego" (quoted from the post i was replying to) (which, as i understand, is a subset of A -- directly from Buckmaster's quote: "the route A is is the route Luis and Diego opened")?

but the post i was replying to claims that "An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and Alpöge’s work", i.e. that the AI model could independently chose B. but choosing B implies choosing A, since B is a subset of A. and in this case it is irrelevant whether Buckmaster claimed that an AI could not independently choose A or B -- the point which you seem to be contesting.

EDIT: my understanding is that "solving the Navier-Stokes existence and smoothness problem" consists of proving at least 1 of 4 precise statements ("options (a) through (d) of Fefferman's statement of the problem"), and Luis and Diego's work were developments towards a proof of statements (c) and (d), which have been recently further expanded by Alpoge and Buckmaster


I just meant that choosing the same 2 routes out of 4 would not at all be a great coincidence without knowing what Tristan was working on.

if you assume the choice of route is a uniformly distributed random variable, yes. but this assumption does not seem consistent with "Almost nobody else I know of was working on it", from Tristan's quote. nor with "It is not the direction one arrives at in a few days by giving a model the problem statement".

It doesn't really make sense what he said. Everyone expects N-S to have blow-up. If you are going to try to show blow-up, C and D are strictly easier than A and B. So Tristan was not clear about what detail of "route" he is talking about, because what he literally said can not be it.

yeah, it looks like i don't have enough knowledge of this subject to discuss this. but i'd be interested in reading what more knowledgeable people have to say about it -- Tristan's work, how different from others' it was (they claim no one else was following the same path), and how it compares with the AI's

Well, no one arrived in this direction, because no one was sitting and prompting a model. It was 10000 agents working 24/7 for several days, trying millions of different directions.

> "It is not the direction one arrives at in a few days by giving a model the problem statement."

Can they back up this statement somehow?


yes, this is ultimately nothing but a claim.

Buckmaster did also mention, for example, that a team of people was employed to solve the problem, which supports this claim. but that is another claim whose veracity could also be questioned. but at some point we must trust other people, unless we can be satisfied with only believing what we personally see.

(also, IMO, the coincidence of both discoveries in time is pretty suspicious. this one doesn't need you to trust many people i guess)


Please don't say brute forced. It sounds like some form of denial or something. Compute for hard problems drops with models--it just means they threw a huge amount of compute. There's (idk about NS specifically so maybe it's exception) no real way to "brute force" a math proof [ok you can enumerate proofs if you can wait until heat death ]

Sorry for random rant but I don't think these statements help your point


I want to point out that almost all previous AI discoveries in math were made in almost the same way. The ideas were there in the community, but weren't considered mainstream/worth pushing forward. Read Tao's comments on the unit distance problem, for example (sry I can't find a link right now).

OpenAI said there [1]: > The method by which the problem was solved is also notable. The proof brings unexpected, sophisticated ideas from algebraic number theory to bear on an elementary geometric question.

[1] https://openai.com/index/model-disproves-discrete-geometry-c...


> And simply knowing a problem can be solved is half the battle.

Have you done any mathematical research? If not, then no, knowing that a problem is solvable is not “half the battle”.

Homework problems are all designed to be solvable, yet they can vary greatly in difficulty. Research mathematics is even more extreme, because, unlike with homework, you don’t know that it is solvable with the extant mathematics, and you might need to invent new maths.


You're taking the phrase too literally. The point is that knowing a solution is possible gives you the conviction to actually find that solution. The hardest part of solving a problem is often a lack of conviction to see it through, and quitting too early. Once you know a solution exists, you can commit maximal effort towards solving it and know that your efforts are not in vain.

If not for the rumors that A/ had already solved NS, OAI would likely never have pursued solving the problem with such fervour. The rumors drove OAI to assemble an entire team to crack this.


How is it unclear? The entire point of deploying models across corporate America is to train on your workflows. Eventually replacing you with digital you is why they're doing it!

> OpenAI looked at user data, stole world class researchers' work

This doesn't seem to be clear and is very implausible for a large company. Be as cynical as you want, but a normal researcher will simply not have access rights to this data, which will be siloed away somewhere else.

It might very well be somewhat unfair to catch wind of a promising approach and then try to frontrun them by throwing compute at the problem, but this isn't really the same.


No, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell

"I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."

Let's see what statement OpenAI will come up with for their side of the story

EDIT: precised my thought on user data vs session data


I was downvoted initially, look what OpenAI shared https://openai.com/index/navier-stokes-solution/ ...

"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve."

"When a further trained version of our internal model became available over the course of the effort"

"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "


If it's siloed the same way the HF bots were, that doesn't exactly bode well. I'd be amazed if there weren't some big companies sending a fleet of lawyers at OpenAI's ZRPs after this news

If you have work happening in a part of your latent space that's got a much lower representation in your dataset then it's pretty plausible to include it. It doesn't actually matter who the user is if there's not a lot of people in the world working on problem X and you have a dataset of work on problem X.

They likely train on logs.

Not only did they look at user data

The whole business is based on reselling user data scraped from the whole internet

It’s plagiarism at scale


Things are more entangled than that. The contribute made from both OpenAI and Anthropic models to solve these problems are clear, now it really hard to quantify which one contributed more, if the role played by the human is major or minor.

OpenAI tried to collaborate and share the results together with a fixed timeline, to avoid this mess but it was inevitable. There is a conflict of interest, where the other researcher works at Anthropic, who will also try to take credit.

Where they may be in the wrong is if they took user data regarding the problem, how will we know if they did or not?


They offered to collaborate by asking to drop a coauthor.

That is not collaboration, and is not an academic norm.


how much of the "codex discussion" was actually ideas also generated by open ai ? this is a bit the elephant in the room

LLMs do nothing but steal, and the companies that own them are fully aware and eager to do it.

Yes, but only if you take this one sided statement at face value.

Why would have they rushed the publication if this was not true? Are you also suggesting that he fully invented the call with Open AI?

The results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.

How else would you explain OpenAI suddendly assembling a team focused on working the same problem from the same angle than the researchers that just made a breakthrough?

What would be so hard to explain?

That OpenAI heard about the result and decided to throw a lot of money at it knowing it was within reach ?

That the model took an approach that was published years ago ?


That's a fake explanation, it skips explaining/justifying how OpenAi "heard about" the result

The rumors that Anthropic had solved a millennium problem were absolutely everywhere last week. I'm not surprised at all that OAI took their own stab at it.

It has no legs as an "explanation" because the content of the email (as described) already acknowledges as much. It's the whole reason he wrote contact email. His chief complaint now includes how the hell did Altman's people know specifically his line of attack down to certain technical keywords. E.g. AI plagiarism.

Where were you seeing those roumors? Care to point to an HN post?

We are told in the statement that there were rumours already going around about a Stokes result (I even know about the rumors in question. It was all over twitter in the right spaces) and that Tristan contacts OpenAI about the rumors. In the message, it's pretty clear Tristan has made some result.

So either the rumors or Tristan's contact would explain it fine.


One should not evaluate explanations based on level of "fine"/innocuousness, in ethics that is called motivated reasoning or something.

The whole point of Tristan's first email was to address the issue of the rumors so in fact this explanation confounds several things (in terms of the mutual knowledge of the conversants and their intentions).

Based on this lack of understanding it is pointless to continue this thread


So are you baselessly assuming that he is lying? He explicitly reported that he was threatened and your answer here is to defend OpenAI no matter what.

Do you not have reading comprehension? Did you even read the statement? He himself asserts at the end he doesn't know if the above example is true or not. What on earth are you going on about? Where in my comment am I assuming he's lying ?

From my post above: > He explicitly reported that *he was threatened* and your answer here is to defend OpenAI no matter what.

Yes, I read fully the statement, what about you? Do you know what is a threat? What is this in your super-humble opinion if not a threat:

> The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

And you are saying that I don't have reading comprehension...


You really lack reading comprehension

My reading comprehension is pretty good, I'm not the one that doesn't recognize a threat even when it's perfectly clear. Verbatim from the statement:

> The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”


I’m open to evidence, but just using Bayesian reasoning, OpenAI is one of the most dishonest companies in history. They’re currently being sued for a dozen employees stealing Apple hardware! I don’t understand why I should give them any grace.

Tailscalar here. Did you expect a wearable fabric or a CNI?


"It can be two things."


late. But was hoping for wearable :)


Tailscalar here.

The Windows client caches the current version for a while, so may not yet have v1.32.3 available on your device. In that case, you can still pull the latest release from http://pkgs.tailscale.com/stable.


Tailscale admin here, politely requesting client update push capability. Being able to see endpoint version is helpful, I will be suspending unpatched endpoints in the near future.


Seconded


Thirded

Or Get Tailscale in the Windows Store so it could auto update for all the endpoints out in the wild I don't control (company laptops, users home PCs, etc). Trying to get TEN people to manually update today was a pain, I can't imagine even triple that.


`winget upgrade tailscale.tailscale` works too.


I think that's right. I would strengthen that statement slightly - it's about ensuring that no actor - whether an insider, or someone who has stolen their credentials, or otherwise compromised them - can perform an action that single handedly accesses user data, without it being known to another actor - via access logs, via approvals, etc.

In terms of the upstream introduction of a new vulnerability, Binary Authorization for Borg can only verify that the code was in fact merged. See the section on third party code, "When importing changes from third party or open source code, we verify that the change is appropriate (for example, the latest version)."

Disclosure: I work at Google and helped write this whitepaper on Binary Authorization for Borg.


I think 'garbage' is a strong word, but I believe what the original poster is trying to say is that there are a lot of binaries, packages, and libraries that most organizations will consume from upstream, and not verify directly. This requires either trust on a third party (often many third parties - in the case of open source), or more intense validation of those components and any changes to those components.

Binary Authorization for Borg performs verification for pieces that come out of Google's CI/CD pipeline. For third party code, see in the doc, "When importing changes from third party or open source code, we verify that the change is appropriate (for example, the latest version)."

Disclosure: I work at Google and helped write this whitepaper on Binary Authorization for Borg.


Yes, there's a few listed in this blog post: https://cloud.google.com/blog/products/identity-security/bey... - Kubernetes admission controllers, OSS part of Kubernetes: https://kubernetes.io/docs/reference/access-authn-authz/admi... - Kritis, OSS: https://opensource.google/projects/kritis - OPA Gatekeeper, OSS: https://github.com/open-policy-agent/gatekeeper - Binary Authorization on GKE/Anthos: https://cloud.google.com/binary-authorization/ They don't all do all the pieces. The hardest part is going to be integrating whatever enforcement solution you choose with your upstream CI/CD pipeline.

Disclosure: I work at Google and helped write this whitepaper on Binary Authorization for Borg.


Agreed with this statement. It's a best practice generally to verify all software updates originate from a particular source before applying them in your environment. Most over the wire updates do this. What's different with Binary Authorization for Borg is that within Google, that last verification step means more than just "came from Google", but "came from Google and went through all previous necessary checks", because of the way the CI/CD system works together.

Disclosure: I work at Google and helped write this whitepaper on Binary Authorization for Borg.


I am perfectly aware that Binary Authorization for Borg is for binaries running inside Google.

I am saying that our solution provides almost the opposite: publicly-verifiable assurance that you are running legitimate binaries, despite being built by automation


in-toto.io also addresses the "proof it went through some steps". How would you compare the two systems?


The biggest difference for me is that in-toto allows you to define any set of upstream metadata requirements, in a very open format, and Binary Authorization has a set of centrally defined requirements, that teams tend to implement in tranches, to meet minimum requirements. It may sound better to have a freeform format, but in practice, I've found that it makes it harder for people to know what they should actually do. In Binary Authorization for Borg, services still define service-specific policies, but pick from a previously defined set of potential requirements. See the section on service-specific policies: https://cloud.google.com/security/binary-authorization-for-b...

You can more easily compare Grafeas and Kritis (OSS projects Google developed, which are similar to GCR Vulnerability Scanning and Binary Authorization for GKE), to in-toto. In fact, I gave a talk covering some of the options for this here: https://youtu.be/uDWXKKEO8NU?t=1314

Disclosure: I work at Google and helped write this whitepaper on Binary Authorization for Borg.


I did a BSc in math/econ, then MSc in math. Went into consulting work, mostly in security. Now work in crypto for a tech company.


Gmail: incl. news alerts

Feedly (fresh articles): Ars Risk Assessment, Bloomberg, The Atlantic Business, various friends' and food blogs

Pocket (older articles)

If I have more time: HN, r/crypto, sometimes Medium

To waste time: Sporcle, Instagram, Foodgawker


Hi, We're the team working on Google's experimental encrypted BigQuery client, also mentioned in the article. It's open-sourced and available on github: https://github.com/google/encrypted-bigquery-client

Happy to answer any questions or chat about cool uses of PHE!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: