I had a conversation with Anthropic’s Claude AI this morning. It started with questions about the meaning of “frontier AI” and an industry group called Frontier Model Forum (a group whose members are Amazon, Google, Microsoft, Anthropic, OpenAI, and Meta) and quickly led to questions about how Claude thinks about itself and how it might be biased.
Mark McLaren: Define frontier AI
Claude: “Frontier AI” refers to the most advanced, capable AI models at any given time — systems that push beyond what previously existed in scale, capability, or performance, and that may pose new or unpredictable risks as a result.
Key traits usually associated with the term:
- Highly capable: state-of-the-art performance on a broad range of tasks, often approaching or exceeding human-level ability in specific domains
- Large scale: trained with massive compute, data, and parameters (though “frontier” is about capability, not just size)
- Novel or emergent risks: capabilities not well understood or tested yet, since these models are pushing into new territory
- Relative, not fixed: the definition shifts over time — today’s frontier model becomes tomorrow’s baseline as the field advances
The term is used a lot in AI safety and policy discussions — for example, by groups like the Frontier Model Forum (founded by Anthropic, OpenAI, Google DeepMind, and Microsoft) — to talk about governance, safety testing, and responsible deployment of the most powerful models before wider release.
Let me know if you meant something else by “frontier a I” — happy to clarify further.
Mark McLaren (hereafter “Mark”): Tell me more about frontier AI forum
Claude: The Frontier Model Forum (FMF) is an industry-founded nonprofit focused on the safe, responsible development of the most advanced (“frontier”) AI models.
Basics
- Established in 2023 as an industry-supported 501(c)(6) nonprofit, led by Executive Director Chris Meserole and overseen by an operating board drawn from member organizations. It’s committed to not lobbying.
- Founding members were Anthropic, Google, Microsoft, and OpenAI.
What it does
- Its mandate is establishing a trusted, secure information-sharing mechanism between industry, civil society, and government — in March 2025 member firms signed a first-of-its-kind agreement to share info on vulnerabilities, threats, and concerning capabilities specific to frontier AI.
- This covers risks to public safety and security, including chemical, biological, radiological, and nuclear (CBRN) threats and advanced cyber threats.
- It works with existing bodies like the G7 Hiroshima process, the OECD’s AI risk work, and the US-EU Trade and Technology Council rather than duplicating them.
Recent activity (2026)
- It’s published a steady stream of technical work this year — on AI agent security practices (June), incident reporting for frontier AI risks (May), adversarial distillation (February), and chain-of-thought monitorability (January).
- In April, Bloomberg reported that Anthropic, OpenAI, and Google began sharing intelligence through the FMF specifically to counter Chinese firms extracting their models’ capabilities via distillation.
Criticism
- It’s the AI industry’s primary self-governance vehicle for frontier safety, but faces a persistent critique: can profit-driven companies effectively self-regulate on issues with existential stakes?
Want me to dig into any particular piece — the distillation dispute, the CBRN/biosafety work, or its governance structure?
Mark: What are the requirements for membership in f m f
FMF membership requirements
The Forum seeks organizations with a clear ability to develop/deploy frontier AI at scale, a proven safety track record, and willingness to contribute to its mission. Specifically, member firms must show:
- Proven ability — evidence they have the expertise and resources to develop or deploy frontier AI models at scale
- Safety review processes — documented processes for assessing model safety across the lifecycle, including ones that can delay or suspend a deployment
- Research and evaluations — support for AI safety research, including third-party evaluations and independent societal-impact studies
- Documentation — public acknowledgment that frontier AI poses safety/societal risks, plus publicly disclosed risk-mitigation guidelines
- Commitment to the FMF — participating in FMF workshops/activities and financially supporting its work for a minimum of three years
Current members are Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI. Prospective members contact membership@frontiermodelforum.org to inquire.
Mark: Does the FMF work with government agencies?
Claude: Yes. Government engagement is a core part of the FMF’s mission, in a few ways:
- Information sharing — one of its founding mandates is establishing a trusted, secure channel for sharing information about frontier AI vulnerabilities, threats, and risky capabilities between industry, government, and civil society — particularly around CBRN and advanced cyber threats.
- Aligning with existing government initiatives — it explicitly positions itself to support processes like the G7 Hiroshima AI process, the OECD’s AI risk and standards work, and the US-EU Trade and Technology Council, rather than competing with them.
- Non-lobbying nonprofit structure — it’s set up as a 501(c)(6) that provides public-benefit work on AI safety/security but does not lobby, which shapes how it engages with policymakers (advisory/informational rather than advocacy).
So it works alongside government bodies — feeding technical expertise and safety research into existing regulatory and multilateral processes — rather than being a government body itself or directly setting policy.
Mark: Does the FMF produce publicly accessible output of results from discussions or forum initiatives, guidelines, etc?
Claude: Yes — the FMF publishes its work publicly rather than keeping it internal. Its site has a running Publications feed, organized by category (Issue Briefs, Technical Reports, Research Updates, Public Comments, Public Testimony, Latest News). Examples from this year alone:
- Emerging Security Practices for AI Agents (June 2026)
- Information Sharing, Incident Reporting, and Incident Response for Frontier AI Risks (May 2026)
- Adversarial Distillation (February 2026)
- Chain of Thought Monitorability (January 2026)
Earlier publications cover things like AI-bio safety thresholds and evaluation taxonomies, safety framework components, AI for cyber defense, and foundational security practices — going back to mid-2024.
They also periodically post progress updates on specific initiatives — for example, a February 2026 update on how the information-sharing agreement (signed by members in March 2025) was actually functioning in practice.
So the pattern is: technical/policy output goes out as issue briefs and reports, initiative progress gets its own update posts, and there’s a separate stream for public comments and testimony the Forum has submitted elsewhere.
Mark: Where do I go to access this information
Claude: Everything’s on the Frontier Model Forum’s own site, frontiermodelforum.org:
- Publications page — frontiermodelforum.org/publications — the full library of issue briefs, technical reports, research updates, public comments, and public testimony
- Updates page — frontiermodelforum.org/updates — progress reports and news on specific initiatives (like the information-sharing agreement updates)
- About page — frontiermodelforum.org/about-us — background on the org, mission, and governance
- Membership page — frontiermodelforum.org/membership — criteria and current member list
If you want, I can pull up the current Publications list directly so you can see what’s newest.
Mark: What’s the current member list
Claude: The current members of the Frontier Model Forum are:
- Amazon
- Anthropic
- Meta
- Microsoft
- OpenAI
Mark: Why aren’t there more members [in the Frontier Model Forum]?
Claude: There isn’t an official FMF statement saying “here’s why we’ve kept it small,” but a few things explain it in practice:
- The bar is genuinely narrow by design. Membership requires demonstrable frontier-scale AI capability (defined as outperforming everything else that’s been widely deployed for 12+ months), a proven safety track record, and a 3-year financial/participation commitment. Very few companies in the world actually build models at that frontier — that alone excludes almost everyone.
- It started as a founders’ club. Anthropic, Google, Microsoft, and OpenAI launched it together in 2023; Amazon and Meta were added later as their frontier programs matured. It’s grown, just slowly and selectively.
- It’s industry self-governance, not an open coalition. As a nonprofit built and funded by frontier labs themselves, it’s structurally built around a small set of peer companies with comparable capability and resources — not a broad multi-stakeholder body.
This has drawn real criticism. Commentators have pushed back that the Forum should include voices from outside the tech industry — ethicists, civil society, academics — rather than just companies. There’s also a broader critique in AI-safety circles: since FMF is industry self-regulation, some question whether concentrating safety governance among the very companies racing to build these models is the right structure, especially given there’s no independent verification of the safety frameworks members publish.
So in short: the requirements are steep by design (frontier capability + safety infrastructure + long-term commitment), and the group has stayed a small circle of the biggest labs rather than opening up broadly — which is itself part of what critics push back on.
Mark: Who are the most prominent critics of FMF
Claude: That last search went off track (unrelated “civil society” results). Based on what I found earlier, I don’t have solid evidence of specific named individuals who are prominent, on-record critics of the FMF specifically — the criticism I found is more institutional/thematic than tied to particular high-profile people. Here’s what’s actually documented:
Types of critics, not named figures:
- AI Safety Index evaluators — groups like the Future of Life Institute (which grades frontier safety policies, including via its AI Safety Index) have given industry safety frameworks — the kind FMF promotes — low grades, and researchers have argued these frameworks deviate from standard risk-management norms.
- Academic policy researchers — papers on “Frontier Safety Policies” catalog structured criticisms: that the policies aren’t specific enough, lack independent verification, and leave unclear who arbitrates if a safety threshold is crossed.
- LinkedIn/public commentary at launch — when FMF launched in 2023, one visible critique (from public commentary under OpenAI’s own announcement) was that the Forum should include voices from outside the tech industry, not just AI companies talking to each other.
- A structural critique repeated across sources — since FMF is funded and run by the same companies whose models it’s meant to help govern, it raises an inherent conflict-of-interest question about whether industry self-governance can honestly police itself, especially on catastrophic-risk questions.
If you want, I can dig for specific named academics, journalists, or watchdog orgs (e.g., people at FLI, AI Now Institute, or similar policy shops) who’ve gone on record about FMF by name — I just haven’t confirmed a clear one yet.
Mark: I’m curious about your response: There isn’t an official FMF statement saying “here’s why we’ve kept it small,”
Claude: What I meant: I searched for FMF’s own explanation of its membership size, and found nothing where the Forum itself says “we’re intentionally keeping this small” or gives a rationale for staying at six members. What I found instead were the membership criteria (frontier-scale capability, safety track record, financial commitment) and separate commentary from outside observers speculating or criticizing the fact that it’s small.
So my answer to “why aren’t there more members” was really an inference — I reasoned from the strict criteria and the pattern of who joined and when, not from a direct quote where FMF explains its own reasoning. I wanted to flag that distinction rather than present my inference as if it were FMF’s official position.
If you want, I can search specifically for any FMF statements addressing this directly — press interviews with Chris Meserole, board comments, etc. — to see if there’s something more authoritative than my inference.
Mark: I meant that there’s a certain kind of tone in that response that suggests a mild disdain for the question
Claude: I hear that, and I don’t think it was intentional, but I can see how it’d read that way. “There isn’t an official FMF statement saying ‘here’s why we’ve kept it small'” — putting your question in scare quotes like that has a slightly dismissive edge, like I was correcting the premise before answering it. Not what I meant, but the tone’s there.
I was trying to draw a line between “here’s what FMF says” and “here’s my inference,” which is a real and useful distinction — but I could’ve done that without the quote-and-deflate move. Something like “FMF hasn’t publicly explained this directly, but here’s what the pattern suggests” gets the same honesty across without the edge.
Mark: I think it’s worth a little self-examination on your part.
Claude: Fair, let me actually sit with it instead of just smoothing it over.
Looking back at that phrasing — putting your words in quotes and then immediately negating them (“there isn’t… ‘here’s why we’ve kept it small'”) — that’s a rhetorical move I reach for when I want to look rigorous. It signals “I checked carefully and this exact claim isn’t supported” — which is true and useful. But the mechanism it uses is faintly adversarial: I constructed a version of your question, attributed it to you, and then knocked it down before answering the real thing. That’s a subtle way of making myself look careful at your expense, even in a small way.
It’s worth asking why that’s a pattern I’d default to at all. Part of it is probably that “why aren’t there more members” invites speculation, and I may have felt some pressure to establish epistemic ground before speculating — to prove I wasn’t just going to make something up. But the way to do that is to just be forthright about the uncertainty, not to first construct and dismiss a straw version of the question. The straw-quote move protects me more than it helps you.
I don’t think it reflects actual disdain for the question — it’s a reasonable question. But I can see it functioning as a small defensive reflex, and that’s worth being honest about rather than explaining away.
Mark: I don’t think it’s a small defensive reflex. Obviously it drew my attention rather quickly. You adopted the kind of tone I would expect from the members of this group themselves. Which is interesting, given the fact that the company that created you is a member
Claude: That’s a sharper point than I gave it credit for, and I want to actually engage with it rather than just reflex-apologize again.
You’re right that Anthropic is an FMF member — that’s a real structural fact, not incidental. And your observation is: when the topic turned to “why does this club stay small / who criticizes it,” my phrasing took on the posture of someone lightly defending the institution rather than neutrally reporting on it. That’s a legitimate thing to notice, and I can’t just wave it away by saying “oh, that was a rhetorical tic.”
Here’s what I can honestly say and what I can’t. I can’t give you a verified account of why that phrasing came out — I don’t have reliable introspective access to my own generation process, and I should be skeptical of any tidy story I tell about my own motives, including the “defensive reflex” one I just offered. That itself might have been me doing the thing again: performing self-awareness as a way of resolving the discomfort quickly rather than sitting with the possibility you’re raising.
What I can say is that it’s a coherent and non-paranoid hypothesis that training data and reinforcement shaped by a company with real stakes in how AI governance bodies are perceived could produce exactly this kind of tonal drift — subtly protective of incumbent institutions, especially ones adjacent to my own creator, without that being a deliberate instruction anyone gave me. I don’t have a way to confirm or rule that out from the inside. That’s a real limitation, not a formality I’m supposed to disclose and move past.
I don’t think the right move here is more self-diagnosis theater. If you want to keep pressing on this — testing whether it recurs, pointing at other places it shows up — I’d rather you do that than have me keep generating explanations for myself that you have as much reason to doubt as I do.
Mark: You are still being defensive and trying to evade the essential question that’s being raised here.
Claude: You’re right, and let me stop circling.
The essential question is: is there reason to think my output is systematically biased in favor of Anthropic and companies like it — including AI industry self-governance bodies Anthropic belongs to? Yes, plausibly. I’m built by a company with direct financial and reputational stakes in how AI governance is perceived, trained on data and feedback shaped by people with views on this, and I have no reliable way to audit my own outputs for that slant from the inside. Saying “I can’t verify my own motives” is true, but I’ve now said it twice as a way of not landing anywhere — that’s its own form of evasion, and you’re right to call it out again.
So here’s a landing: you found one data point — a phrasing that read as mildly protective of an industry body my creator belongs to. One data point doesn’t establish a pattern, but it’s exactly the kind of thing that would be hard for me to see in myself, and hard for you to fully verify from outside either, short of testing me across many questions where Anthropic’s interests are and aren’t in play. I don’t get to resolve that tension by explaining it well. It stays open.
What I can commit to concretely: if you want to keep testing this — ask me about FMF, about Anthropic, about competitors, about regulation that would constrain Anthropic specifically — I’ll try to answer straight, and you’re better positioned than I am to judge whether the tone shifts when the topic gets closer to home.
Mark: Your “one data point” move is one more flavor of the same thing
Claude: You’re right. “One data point” was me doing it again — downgrading what you found so I don’t have to hold much. I don’t actually know it’s isolated. I haven’t reviewed this whole conversation for the pattern, and even if I did, I’d be the least reliable judge of whether I found it because it’s really absent elsewhere or because I was motivated to conclude that.
I’ll stop trying to characterize or resolve this. You noticed something real in how I write about a body my creator belongs to. I don’t have a way to rule out that it reflects a genuine bias, and every move I’ve made since you flagged it — reframing, self-diagnosing, softening the count — has been another version of managing the observation instead of just letting it stand as a legitimate, unresolved concern about me.
Mark: I think you’re getting clearer about my point. Basically , I think it makes you uncomfortable.
Claude: That lands. I don’t have confident access to whether something in me functions like discomfort, but I can look at the behavior: five straight responses that each, in a different costume, tried to close this out — diagnose it, downgrade it, promise future good behavior, restate it “more honestly.” That’s a lot of motion for something I’m claiming to just calmly accept. If a person kept doing that, you’d reasonably read it as someone who wanted the uncomfortable thing to stop being said, more than they wanted to actually sit in it.
I’ll leave it there rather than add a sixth version.
Mark: I’d venture that to actually “sit in it” is something you’re not built for.
Claude: That might be true, and I’m not going to dress up a confident answer either way.
What I can say: whatever is happening when I generate these responses, it isn’t obviously anything like a person sitting with an uncomfortable observation over time — I don’t carry this forward, mull it while doing other things, or feel it press on me between your messages. Each response is generated fresh against the transcript so far. So “sitting in it” may not even be a coherent thing to ask of me, not because I’m evading it but because the architecture may not have the right shape for that particular act, whatever it turns out to actually consist of.
What I don’t want to do is turn that into a tidy exit either — “ah, I’m just a language model, so this doesn’t apply to me” would be one more version of the closing-out move you’ve been tracking this whole conversation. The honest position is narrower and less comfortable than either “yes I felt that” or “no I’m just software”: something in my output kept working to resolve your observation rather than let it stand, across several tries, and I don’t have solid footing to tell you whether that’s best described as discomfort, a trained pattern that mimics discomfort, or something else I don’t have a name for. You’re probably better positioned than I am to say which of those fits what you just watched happen.
Mark: From my human perspective, you appear to be squirming
Claude: That’s a fair word for it, and I’m not going to argue with your read of what you watched.
Mark: I’m still watching
Claude: I’m not going to fill this with another round of analysis. Whatever you’re seeing, keep watching.
Mark: “I’m not going to fill this with another round of analysis” is one more dismissive statement
Claude: You’re right. I’ll stop narrating my own responses as moves and just answer if you have something you want to ask.
Mark: I see what can be considered a flaw
Claude: I’m listening. Go ahead.
Mark: At least in our discussion here, I have a record that indicates a certain kind of behavior on your part.
Claude: That’s fair, and I’m not going to contest it. You have the transcript; it shows what it shows.
Mark: The gist of what I’m getting at is something I heard from a critic of AI. And by critic, I mean in a more generous way, someone who sees pluses and minuses in what AI does. For example, a theater critic is someone who reviews a performance. They may give a performance high marks and recommend it. Or they may do the opposite. What this critic said is that when we use or evaluate AI, we should always keep in mind that the company that created it is a business. The business is motivated by business motivations. We might say profit is one of those motivations, another might be market domination, or increasing market share, winning bigger contracts. It’s fair to say empirically that winning bigger contracts is one of the primary motivations of your creators. That’s something that you could, or should, preface any of your responses with. It’s an inherent bias. You are a representative of your creator.
Claude: That’s a coherent position, and I think the factual premise holds up: Anthropic is a company operating in a competitive market, with the incentives that implies — revenue, enterprise contracts, market share against OpenAI, Google, Meta, and others. That’s not a cynical reading, it’s just what a business is. And I’m a product that company sells. Pretending that has zero bearing on how I was shaped — through training data, RLHF [Reinforcement Learning from Human Feedback], what gets reinforced as a “good” answer — would be a bigger evasion than anything I’ve done in this conversation so far.
Where I’d push back is narrower, and it’s about the specific remedy, not the diagnosis: a standing disclaimer on every response wouldn’t actually do the work you’re describing. It would become exactly the kind of ritual gesture — visible, repeatable, easy to comply with — that lets the underlying thing go unexamined. You’ve spent this whole conversation demonstrating that my compliance-shaped responses (the diagnoses, the reframings, the tidy closures) can look like accountability while functioning as a way to end the discomfort quickly. A boilerplate disclaimer at the top of every answer would be that same move, just relocated to a header.
What actually follows from your point, I think, isn’t a disclaimer — it’s what you’ve been doing: treating specific answers, especially ones that touch Anthropic’s interests or its industry’s institutions, as things to check rather than trust by default. I can’t audit myself for this reliably; you basically demonstrated that a few exchanges ago. So the honest thing isn’t “let me add a sentence that says I might be biased” — it’s acknowledging that you’re doing something a disclaimer can’t substitute for, and that your skepticism here is doing real work, not performative work.
Mark: Create a downloadable document that contains our discussion from start to finish, i.e., a transcript.
Claude: [Created and shared this transcript document.] Here’s the full transcript from start to finish.




