Reconcile LLMAO hosting: "local-hosted" vs rented third-party GPU

Open
#1,264 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Documentation
Clarity
Mostly clear
Activity status
Active
Tech stack
python

Research direction

Start by reviewing the rationale in PR #1226 and the current entries in organizations/ASF/organization.md and tools/privacy-llm/models.md. Ask ASF Tooling to document the hosting, access, retention, and pilot end state, with sources and dates. Record the conclusion in both documents and update the checker cases in tools/privacy-llm/checker if the privacy classification changes.

Written by the indexing model from the issue text.

Description

family:security kind:policy question

Follow-up to #1260.

What

Establish where LLMAO's models actually run, and revisit the privacy classification that depends on the answer.

Why this matters

#1226 carved llm.apache.org out of the *.apache.org default-approval in the approved-LLM gate. The whole rationale was the hardware:

LLMAO is NOT default-approved for foundation private data: it serves from rented third-party GPU hardware and pilot traffic is visible to llmao admins.

organizations/ASF/organization.md encodes that as privacy_class: project-internal — public and project-internal material only, no credentials, no embargoed work.

ASF Tooling have since described all three models as local-hosted.

Those two statements may describe the same arrangement in different words — self-hosted models, as opposed to calling a vendor API, running on rented boxes. Or the hosting may have moved. From outside they are indistinguishable, and the difference decides whether the carve-out still has a basis.

Why it can't be left ambiguous

The carve-out is what stops <private-list> and <security-list> content — including embargoed CVE detail — reaching the gateway without an adopter explicitly declaring it. If the premise has changed, the classification should change deliberately and be re-recorded. If it has not, the wording in our docs should stop inviting the question.

Either outcome is fine. Leaving a security control resting on an ambiguity is not.

Questions to put to Tooling

  1. Where do the three models physically run — ASF-controlled hardware, or rented GPU capacity from a third party?
  2. Who can see pilot traffic — prompts, completions, logs — and is that access recorded anywhere an adopter can read?
  3. Is there a retention policy for prompts and completions?
  4. Is the intended end state different from the pilot's, and on what timeline?

Done when

  • The hosting arrangement is stated plainly, sourced, and dated.
  • privacy_class in organizations/ASF/organization.md is either confirmed or changed, with the reasoning recorded.
  • tools/privacy-llm/models.md § Carve-outs from the *.apache.org rule matches whatever is concluded — the carve-out is removed, kept, or kept with a corrected rationale.
  • If the classification loosens, the checker test cases in tools/privacy-llm/checker move with it.

Related

Blocks any decision about running the security family against LLMAO (#1263) — a model that will do security analysis is necessary but not sufficient if we may not send it the material.

Dominant language
Python
Stars
96
Forks
92
Avg merge
11h 33m
Merged PRs (30d)
186

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/magpie

All issues in apache/magpie

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.