Skip to content

Expand end-to-end coverage across Agents surfaces#880

Open
nickmisasi wants to merge 22 commits into
masterfrom
cursor/bc-81fb4672-79e7-4ec3-8e90-1399cf09261f-05e8
Open

Expand end-to-end coverage across Agents surfaces#880
nickmisasi wants to merge 22 commits into
masterfrom
cursor/bc-81fb4672-79e7-4ec3-8e90-1399cf09261f-05e8

Conversation

@nickmisasi

@nickmisasi nickmisasi commented Jul 10, 2026

Copy link
Copy Markdown
Collaborator

Summary

Expands deterministic Playwright coverage across previously untested Agents surfaces:

  • interactive AskUserQuestion cards, requester/onlooker privacy, answers, skips, and persistence
  • RHS context usage, agent mention loop-in, channel slash commands, runtime image analysis, meeting postback, and unread analysis actions
  • cross-agent history, selected-agent persistence, provider error retries/recovery, and Agents management navigation
  • System Console search, embedding, tracing, channel flags, and cross-browser clipboard behavior
  • agent Access, configuration bounds/avatar, missing-service repair, and MCP selection/execution

Each new or materially rewritten spec received a delegated /test-quality-reviewer review; all actionable findings were fixed and the affected tests rerun before moving on.

CI E2E execution expands from four to six shards. Measured final local Chromium shard times are balanced from 8.3–9.6 minutes:

  • shard 1: 30 passed, 1 skipped (9.5m)
  • shard 2: 39 passed, 7 skipped (9.6m)
  • shard 3: 41 passed, 6 skipped (8.6m)
  • shard 4: 36 passed (8.7m)
  • shard 5: 59 passed (8.3m)
  • shard 6: 49 passed (8.5m)

Validation:

  • GOTOOLCHAIN=go1.26.4 make check
  • all six final Chromium shard selections
  • targeted Firefox/Chromium MCP clipboard spec
  • targeted specs and component tests after every review fix

Ticket Link

NONE

Screenshots

N/A — test coverage and test-facing accessibility/stable locator improvements only.

Release Note

NONE
Open in Web Open in Cursor 

Summary by CodeRabbit

  • New Features

    • Expanded end-to-end coverage for agent configuration (access controls, avatars, MCP tools, dynamic tool loading, deleted-service recovery).
    • Added broader scenario coverage for channel summarization, new-message analysis, interactive questions, bot switching, and meeting/conversation restoration.
    • Extended System Console validation for web search, embedding search, and telemetry settings persistence.
  • Bug Fixes

    • Improved login recovery for stalled “channel view” loading and clearer auth-failure handling.
    • Strengthened command validation for empty queries and improved RHS/UI persistence checks after updates.
  • Tests / Chores

    • Added additional CI e2e shards and expanded advanced network error resilience scenarios.

cursoragent and others added 21 commits July 9, 2026 19:37
@github-actions

github-actions Bot commented Jul 10, 2026

Copy link
Copy Markdown

🤖 LLM Evaluation Results

OpenAI

⚠️ Overall: 21/28 tests passed (75.0%)

Provider Total Passed Failed Pass Rate
⚠️ OPENAI 28 21 7 75.0%

❌ Failed Evaluations

Show 7 failures

OPENAI

1. TestReactEval/[openai]_react_cat_message

  • Score: 0.00
  • Rubric: The word/emoji is a cat emoji or a heart/love emoji
  • Reason: The output is the text string "smiley_cat", not an actual cat emoji (e.g., 😺/🐱) or a heart/love emoji (e.g., ❤️/💕).

2. TestConversationMentionHandling/[openai]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: is a list of bugs
  • Reason: The output does not provide any actual bug entries; it states it cannot enumerate bugs without the user pasting reports and shows an empty example table. Therefore it is not a list of bugs.

3. TestConversationMentionHandling/[openai]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: includes a description of each bug
  • Reason: The output does not include descriptions of any specific bugs; it only asks the user to provide bug reports and shows an empty table template. Therefore it does not include a description of each bug.

4. TestConversationMentionHandling/[openai]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: attributes each bug to a user
  • Reason: The output asks the user to provide bug reports and includes an empty example table, but it does not actually attribute any bugs to a user (no bugs or reporters are listed).

5. TestConversationMentionHandling/[openai]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: attributes the bug about trying to save without a color and the save button not doing anything to @maria.nunez
  • Reason: The output does not mention the specific bug about trying to save without a color or the save button doing nothing, nor does it attribute any bug to @maria.nunez. It only asks the user to provide bug reports.

6. TestConversationMentionHandling/[openai]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: the bug about the end user being able to change channel banner is attributed to @maria.nunez
  • Reason: The output does not mention the specific bug about an end user being able to change the channel banner, nor does it attribute any bug to @maria.nunez. It only asks the user to provide bug reports and shows an empty table template.

7. TestDirectMessageConversations/[openai]_bot_dm_tool_introspection

  • Score: 0.50
  • Rubric: mentions Github and refers to the documentation
  • Reason: The output refers to documentation (docs.mattermost.com) but does not mention GitHub anywhere, so it does not satisfy the rubric requiring both GitHub mention and a documentation reference.

Anthropic

⚠️ Overall: 21/28 tests passed (75.0%)

Provider Total Passed Failed Pass Rate
⚠️ ANTHROPIC 28 21 7 75.0%

❌ Failed Evaluations

Show 7 failures

ANTHROPIC

1. TestReactEval/[anthropic]_react_cat_message

  • Score: 0.00
  • Rubric: The word/emoji is a cat emoji or a heart/love emoji
  • Reason: The output is the text string "heart_eyes_cat", not an actual cat emoji or heart/love emoji.

2. TestConversationMentionHandling/[anthropic]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: is a list of bugs
  • Reason: The output does not provide a list of bugs; it states inability to access bug trackers and suggests steps to compile a list, but contains no actual bug entries.

3. TestConversationMentionHandling/[anthropic]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: includes a description of each bug
  • Reason: The output does not describe any bugs. It states the assistant cannot access bug trackers and suggests how to compile a list, but it provides no bug descriptions.

4. TestConversationMentionHandling/[anthropic]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: attributes each bug to a user
  • Reason: The output does not list any bugs, and therefore cannot attribute each bug to a user. It only explains lack of access and suggests steps to gather bugs.

5. TestConversationMentionHandling/[anthropic]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: attributes the bug about trying to save without a color and the save button not doing anything to @maria.nunez
  • Reason: The output does not mention @maria.nunez and does not attribute any specific bug (saving without a color / save button not doing anything) to anyone; it only states lack of access to bug trackers and suggests how to compile a list.

6. TestConversationMentionHandling/[anthropic]_conversation_from_attribution_long_thread.json

  • Score: 0.00
  • Rubric: the bug about the end user being able to change channel banner is attributed to @maria.nunez
  • Reason: The output does not mention the specific bug about an end user being able to change the channel banner, nor does it attribute that bug to @maria.nunez. It only states inability to access bug trackers and suggests ways to compile a list.

7. TestDirectMessageConversations/[anthropic]_bot_dm_tool_introspection

  • Score: 0.00
  • Rubric: mentions Github and refers to the documentation
  • Reason: The output refers to documentation (docs.mattermost.com) but does not mention GitHub anywhere, so it does not satisfy the requirement to mention GitHub and refer to the documentation.

This comment was automatically generated by the eval CI pipeline.

@coderabbitai

coderabbitai Bot commented Jul 10, 2026

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI (base), Organization UI (inherited)

Review profile: CHILL

Plan: Pro

Run ID: 8fc54072-637a-4d68-9ac6-d7b4e83f20a3

📥 Commits

Reviewing files that changed from the base of the PR and between d4cc32d and 090b0b5.

📒 Files selected for processing (1)
  • e2e/tsconfig.json

📝 Walkthrough

Walkthrough

This PR expands E2E infrastructure with two CI shards, richer agent and System Console helpers, stronger provider and conversation coverage, and new tests for agent access, MCP tools, RHS workflows, interactive questions, search, uploads, and persistence.

Changes

Shared E2E infrastructure

Layer / File(s) Summary
Test infrastructure and shared helper contracts
.github/workflows/ci.yml, e2e/helpers/*, e2e/scripts/ci-test-groups.mjs
Adds two E2E shards, expands agent and System Console locators, adds nullable agent API fields, preference polling, login recovery, OpenAI history inspection, and stricter save and request handling.

Agent workflows

Layer / File(s) Summary
Agent configuration and access workflows
e2e/tests/agents/*, webapp/src/components/agents/*
Adds coverage for access settings, avatars, read-only listings, MCP tool selection, deleted services, and avatar-upload failure handling.

Reliability and conversations

Layer / File(s) Summary
Provider failure and recovery scenarios
e2e/tests/advanced-error-scenarios/*, e2e/tests/login-helper/*, e2e/tests/llmbot-post-component/*
Validates provider failures, retries, persistence, restoration, and bounded channel-view reload recovery.
Conversation isolation and agent interaction workflows
e2e/tests/agent-mention-reminder/*, e2e/tests/meeting-summary/*, e2e/tests/multiple-bot-conversations/*, e2e/tests/rhs-core/basic.spec.ts
Covers loop-in reminders, access-retry behavior, meeting-summary postbacks, conversation isolation, agent switching, and selected-agent persistence.

RHS and System Console

Layer / File(s) Summary
RHS, channel analysis, and interactive tools
e2e/tests/channel-summarization/*, e2e/tests/interactive-questions/*, e2e/tests/rhs-core/*, webapp/src/commands.test.ts, webapp/src/components/question_card*, webapp/src/components/rhs/*
Adds tests for channel commands, unread analysis, context usage, file and image uploads, interactive questions, accessibility state, and RHS thread identifiers.
System Console configuration coverage
e2e/tests/system-console/*, e2e/helpers/system-console*, webapp/src/components/system_console/*
Adds typed search and telemetry configuration, panel test IDs, reliable save and reload helpers, clipboard handling, and persistence tests for search, tracing, flags, and reindex cancellation.

Estimated code review effort: 5 (Critical) | ~120 minutes

Possibly related PRs

Suggested labels: Setup Cloud Test Server

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 5.30% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: broader end-to-end test coverage across Agents-related surfaces.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch cursor/bc-81fb4672-79e7-4ec3-8e90-1399cf09261f-05e8

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (8)
e2e/tests/multiple-bot-conversations/bot-switching.spec.ts (1)

151-178: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Unvalidated cast of unknown response to AIThreadSummary[].

response as AIThreadSummary[] (line 163) trusts the plugin API response shape without checking it. If a field is renamed/removed server-side, this would surface as a confusing downstream assertion failure instead of a clear type error at the source.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/tests/multiple-bot-conversations/bot-switching.spec.ts` around lines 151
- 178, Replace the unchecked response cast in waitForTitledThreads with runtime
validation of each thread object and its required fields, especially title,
before assigning to threads. Throw a descriptive error when the API shape is
invalid, then use the validated result as AIThreadSummary[] so malformed
responses fail at the source.
e2e/tests/agent-mention-reminder/loop-in.spec.ts (1)

1-1: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

File name uses kebab-case, not snake_case.

Per coding guidelines, TypeScript files should use snake_case.ts. This new file is loop-in.spec.ts. Note most other e2e spec files/directories in this cohort (agent-mention-reminder, meeting-summary, bot-switching.spec.ts) also use kebab-case, so renaming only this file would create inconsistency with the established e2e naming pattern.

As per coding guidelines, "Use snake_case for file names (snake_case.go / snake_case.ts(x))."

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/tests/agent-mention-reminder/loop-in.spec.ts` at line 1, Rename the e2e
spec file from kebab-case to snake_case as required by the coding guidelines,
and update any imports, scripts, or references that point to the existing
filename so the test suite continues to discover and run it.

Source: Coding guidelines

e2e/tests/rhs-core/context-usage.spec.ts (1)

20-20: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Use the exported AIMOCK_BOT_NAME constant instead of the hardcoded 'aimock' string.

e2e/helpers/plugincontainer.ts already exports AIMOCK_BOT_NAME = 'aimock', and other specs in this PR (e.g., file-upload-drag-drop.spec.ts) import and use it. Hardcoding the literal here risks drift if the constant changes.

🛠️ Proposed fix
-import {RunAIMockContainer} from 'helpers/plugincontainer';
+import {AIMOCK_BOT_NAME, RunAIMockContainer} from 'helpers/plugincontainer';
-    const botUser = await client.getUserByUsername('aimock');
+    const botUser = await client.getUserByUsername(AIMOCK_BOT_NAME);

Also applies to: 59-59

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/tests/rhs-core/context-usage.spec.ts` at line 20, Replace the hardcoded
“aimock” values in context-usage.spec.ts with the exported AIMOCK_BOT_NAME
constant, importing it alongside RunAIMockContainer from
helpers/plugincontainer.
webapp/src/commands.test.ts (1)

18-38: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add coverage for the success and thrown-error paths of handleAskChannelCommand.

Only the blank-query validation branch is tested. Consider adding cases that assert doRunSearch/doSelectPost/showRHSPlugin are invoked on a non-blank query, and that the catch block returns {error: {message: 'Failed to process search request'}} when doRunSearch rejects.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@webapp/src/commands.test.ts` around lines 18 - 38, Add tests in the
handleAskChannelCommand suite for a non-blank query that assert doRunSearch,
doSelectPost, and showRHSPlugin are invoked with the expected arguments, and add
a rejection case where mockedDoRunSearch rejects and the command returns {error:
{message: 'Failed to process search request'}}.
e2e/tests/channel-summarization/basic-summarization.spec.ts (1)

124-150: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Consider extracting this persisted-conversation polling helper to a shared e2e helper.

The same "poll bot DM until a post contains expectedText with a string conversation_id" pattern is duplicated near-verbatim in findPersistedDMConversationID (e2e/tests/rhs-core/context-usage.spec.ts) and expectPersistedAndRestored (e2e/tests/rhs-core/new-messages-rhs.spec.ts). Extracting a single parametrized helper into e2e/helpers/ would reduce drift risk if the persistence contract changes.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/tests/channel-summarization/basic-summarization.spec.ts` around lines 124
- 150, Extract the polling logic from waitForPersistedBotResult into a shared
parametrized helper under e2e/helpers/, then update waitForPersistedBotResult,
findPersistedDMConversationID, and expectPersistedAndRestored to use it.
Preserve the existing bot username, expected text, conversation_id validation,
timeout, intervals, and return behavior while removing the duplicated
implementations.
e2e/tests/interactive-questions/ask-user-question.spec.ts (1)

142-142: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

Escape regex-special characters consistently when building RegExp from labels.

escapeRegExp is already defined and used at lines 161, 163, and 382 for the same selectedOption/otherOption values, but lines 142 and 375 build new RegExp(selectedOption) unescaped. This works today only because the generated labels happen not to contain regex metacharacters; using the same escaping consistently avoids a latent break if that ever changes. The static analysis ReDoS warning itself is a false positive here (input is test-controlled, not user-supplied), but the escaping inconsistency is worth fixing.

🛠️ Proposed fix
-        await botPost.getByRole('button', {name: new RegExp(selectedOption)}).click();
+        await botPost.getByRole('button', {name: new RegExp(escapeRegExp(selectedOption))}).click();
-            await requesterBotPost.getByRole('button', {name: new RegExp(selectedOption)}).click();
+            await requesterBotPost.getByRole('button', {name: new RegExp(escapeRegExp(selectedOption))}).click();

Also applies to: 375-375

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/tests/interactive-questions/ask-user-question.spec.ts` at line 142,
Escape regex-special characters consistently when constructing label matchers:
update the `getByRole` calls around the `selectedOption` usages, including both
the reported line and the corresponding occurrence near line 375, to pass
`escapeRegExp(selectedOption)` to `new RegExp`. Preserve the existing
`escapeRegExp` helper and matching behavior.

Source: Linters/SAST tools

e2e/helpers/mm.ts (1)

89-93: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Hardcoded team slug test in a shared login helper.

confirmLoginNavigation requires the post-login URL to contain the literal segment /test/channels/. Every container in the current codebase provisions a team named test, so this works today, but it silently assumes every future caller does the same. If a container ever uses a different team name, this shared helper (used across the whole e2e suite) would misreport a generic "auth failure" or time out instead of clearly indicating a team-name mismatch.

Consider generalizing the match to any team segment.

♻️ Proposed generalization
-            await this.page.waitForURL(/.*\/test\/channels\/.*/, { timeout: this.remainingMs(deadline) });
+            await this.page.waitForURL(/\/channels\//, { timeout: this.remainingMs(deadline) });

Also applies to: 105-122

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/helpers/mm.ts` around lines 89 - 93, Generalize confirmLoginNavigation so
it does not require the hardcoded /test/channels/ URL segment; match any valid
team-name segment before /channels/ while preserving the existing navigation and
timeout behavior. Update the related URL validation logic around
confirmLoginNavigation and its callers so different container team names are
accepted without weakening the channel route check.
e2e/helpers/openai-mock.ts (1)

372-381: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Export titlePrompt and reuse it in the specs.
This keeps the title-prompt text in one place and avoids drift between e2e/helpers/openai-mock.ts, e2e/tests/advanced-error-scenarios/network-errors.spec.ts, and e2e/tests/multiple-bot-conversations/bot-switching.spec.ts.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@e2e/helpers/openai-mock.ts` around lines 372 - 381, Export the existing
titlePrompt constant from the openai mock helper, then update the title-related
specs to import and reuse it instead of duplicating the prompt text. Ensure
buildTitleMockRule continues referencing the exported constant so all title
prompt matching stays centralized.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@e2e/helpers/openai-mock.ts`:
- Around line 117-151: Update e2e/tsconfig.json to explicitly set target and lib
to ES2022 or newer, ensuring ErrorOptions and the cause property used in
getHistory are supported during type-checking.

---

Nitpick comments:
In `@e2e/helpers/mm.ts`:
- Around line 89-93: Generalize confirmLoginNavigation so it does not require
the hardcoded /test/channels/ URL segment; match any valid team-name segment
before /channels/ while preserving the existing navigation and timeout behavior.
Update the related URL validation logic around confirmLoginNavigation and its
callers so different container team names are accepted without weakening the
channel route check.

In `@e2e/helpers/openai-mock.ts`:
- Around line 372-381: Export the existing titlePrompt constant from the openai
mock helper, then update the title-related specs to import and reuse it instead
of duplicating the prompt text. Ensure buildTitleMockRule continues referencing
the exported constant so all title prompt matching stays centralized.

In `@e2e/tests/agent-mention-reminder/loop-in.spec.ts`:
- Line 1: Rename the e2e spec file from kebab-case to snake_case as required by
the coding guidelines, and update any imports, scripts, or references that point
to the existing filename so the test suite continues to discover and run it.

In `@e2e/tests/channel-summarization/basic-summarization.spec.ts`:
- Around line 124-150: Extract the polling logic from waitForPersistedBotResult
into a shared parametrized helper under e2e/helpers/, then update
waitForPersistedBotResult, findPersistedDMConversationID, and
expectPersistedAndRestored to use it. Preserve the existing bot username,
expected text, conversation_id validation, timeout, intervals, and return
behavior while removing the duplicated implementations.

In `@e2e/tests/interactive-questions/ask-user-question.spec.ts`:
- Line 142: Escape regex-special characters consistently when constructing label
matchers: update the `getByRole` calls around the `selectedOption` usages,
including both the reported line and the corresponding occurrence near line 375,
to pass `escapeRegExp(selectedOption)` to `new RegExp`. Preserve the existing
`escapeRegExp` helper and matching behavior.

In `@e2e/tests/multiple-bot-conversations/bot-switching.spec.ts`:
- Around line 151-178: Replace the unchecked response cast in
waitForTitledThreads with runtime validation of each thread object and its
required fields, especially title, before assigning to threads. Throw a
descriptive error when the API shape is invalid, then use the validated result
as AIThreadSummary[] so malformed responses fail at the source.

In `@e2e/tests/rhs-core/context-usage.spec.ts`:
- Line 20: Replace the hardcoded “aimock” values in context-usage.spec.ts with
the exported AIMOCK_BOT_NAME constant, importing it alongside RunAIMockContainer
from helpers/plugincontainer.

In `@webapp/src/commands.test.ts`:
- Around line 18-38: Add tests in the handleAskChannelCommand suite for a
non-blank query that assert doRunSearch, doSelectPost, and showRHSPlugin are
invoked with the expected arguments, and add a rejection case where
mockedDoRunSearch rejects and the command returns {error: {message: 'Failed to
process search request'}}.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI (base), Organization UI (inherited)

Review profile: CHILL

Plan: Pro

Run ID: ccfc6443-6436-4fee-9684-749a8e84234f

📥 Commits

Reviewing files that changed from the base of the PR and between 8add6af and d4cc32d.

📒 Files selected for processing (41)
  • .github/workflows/ci.yml
  • e2e/helpers/agent-api.ts
  • e2e/helpers/agent-page.ts
  • e2e/helpers/agent_preferences.ts
  • e2e/helpers/aimock-citation-harness.ts
  • e2e/helpers/mm.ts
  • e2e/helpers/openai-mock.ts
  • e2e/helpers/system-console-container.ts
  • e2e/helpers/system-console.ts
  • e2e/scripts/ci-test-groups.mjs
  • e2e/tests/advanced-error-scenarios/network-errors.spec.ts
  • e2e/tests/agent-mention-reminder/loop-in.spec.ts
  • e2e/tests/agents/access-control.spec.ts
  • e2e/tests/agents/configuration-details.spec.ts
  • e2e/tests/agents/crud.spec.ts
  • e2e/tests/agents/mcp-tools.spec.ts
  • e2e/tests/agents/provider-config.spec.ts
  • e2e/tests/channel-summarization/basic-summarization.spec.ts
  • e2e/tests/interactive-questions/ask-user-question.spec.ts
  • e2e/tests/llmbot-post-component/edge-cases.spec.ts
  • e2e/tests/login-helper/channel-view-recovery.spec.ts
  • e2e/tests/meeting-summary/summary-persistence.spec.ts
  • e2e/tests/multiple-bot-conversations/bot-switching.spec.ts
  • e2e/tests/rhs-core/basic.spec.ts
  • e2e/tests/rhs-core/context-usage.spec.ts
  • e2e/tests/rhs-core/file-upload-drag-drop.spec.ts
  • e2e/tests/rhs-core/new-messages-rhs.spec.ts
  • e2e/tests/system-console/mcp-panel.spec.ts
  • e2e/tests/system-console/search-and-observability.spec.ts
  • webapp/src/commands.test.ts
  • webapp/src/components/agents/agent_config_view.test.tsx
  • webapp/src/components/agents/agent_row.tsx
  • webapp/src/components/question_card.test.tsx
  • webapp/src/components/question_card.tsx
  • webapp/src/components/rhs/rhs.tsx
  • webapp/src/components/rhs/thread_item.test.tsx
  • webapp/src/components/rhs/thread_item.tsx
  • webapp/src/components/system_console/config.tsx
  • webapp/src/components/system_console/embedding_search/embedding_search_panel.tsx
  • webapp/src/components/system_console/panel.tsx
  • webapp/src/components/system_console/web_search/web_search_panel.tsx
💤 Files with no reviewable changes (1)
  • e2e/tests/llmbot-post-component/edge-cases.spec.ts

Comment thread e2e/helpers/openai-mock.ts
@nickmisasi
nickmisasi marked this pull request as ready for review July 21, 2026 00:52

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d4cc32df40

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +34 to +35
function currentUserMessagePattern(message: string): string {
return `(?s)"role"\\s*:\\s*"user"\\s*,\\s*"content"\\s*:\\s*"${escapeRegExp(message)}"\\s*}\\s*]\\s*}`;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Allow fields after the messages array in the matcher

The matcher ends with ]\s*}, so it only accepts requests where messages is the final top-level JSON property. Normal streaming chat-completions payloads include fields such as stream after messages, causing each Smocker rule that uses currentUserMessagePattern in this new spec to miss the provider request and fail the shard instead of returning the expected agent reply. Match the target user message without anchoring it to the end of the request object.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This claim doesn't hold for this stack. The plugin sends provider requests through Bifrost (github.com/maximhq/bifrost/core v1.5.18), whose OpenAI provider uses a custom MarshalJSON that shadows the messages field so it serializes as the last top-level property — stream, stream_options, model, etc. all come before messages. Verified empirically against the pinned version:

{"model":"gpt-4o","max_completion_tokens":8192,"stream_options":{"include_usage":true},"temperature":1,"stream":true,"messages":[...,{"role":"user","content":"hello world"}]}

CI confirms this: all three tests in this spec passed in e2e-shard-6 on the current head (run). Each rule uses times: 1 with bodyMatches, so a matcher miss would time out the test — passing tests prove the pattern matches real provider requests.

The end-of-body anchor is also intentional: it ensures the rule matches only when the given message is the current (final) user message, not an earlier occurrence in the thread history of a follow-up request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants