fix(tests): align CI tests with recent security hardening (#31470)
Four recent security PRs landed on main with stale/missing test updates, breaking 4 test shards on every subsequent PR's CI run: - test_discord_bot_auth_bypass.py (PR #30742c3caca658): DISCORD_ALLOWED_ROLES no longer bypasses _is_user_authorized. Inverted 3 tests to assert the new (correct) behavior: role config alone does NOT authorize at the gateway layer. - test_msgraph_webhook.py (PR #301694ca77f105): adapter.is_connected is a @property, not a method. Test was calling it with () after the connect() change; TypeError: 'bool' is not callable. Removed the parens. - test_feishu_approval_buttons.py (PR #30744bdb97b857): Card-action callbacks now go through _allow_group_message authorization. 3 tests in TestCardActionCallbackResponse didn't populate adapter._allowed_group_users so the operator's open_id got rejected. Added the allowlist setup to each test, matching the existing pattern in test_returns_card_for_approve_action. Also raise tolerance on test_wait_for_process_kills_subprocess_on_keyboardinterrupt: the SIGTERM → 3s TimeoutStopSec → SIGKILL → reap chain can exceed 10s under loaded xdist (40 workers). Bumped _wait_for_pgid_exit timeout 10→30s and worker join timeout 5→15s. Passes 100% in isolation already; this just makes it tolerant of CI-host load. Validation: 270/270 tests pass across the 5 affected files.
This commit is contained in:
@@ -48,8 +48,14 @@ def _process_group_snapshot(pgid: int) -> str:
|
||||
).stdout.strip()
|
||||
|
||||
|
||||
def _wait_for_pgid_exit(pgid: int, timeout: float = 10.0) -> bool:
|
||||
"""Wait for a process group to disappear under loaded xdist hosts."""
|
||||
def _wait_for_pgid_exit(pgid: int, timeout: float = 30.0) -> bool:
|
||||
"""Wait for a process group to disappear under loaded xdist hosts.
|
||||
|
||||
The cleanup chain is: SIGTERM → 3s TimeoutStopSec → SIGKILL → reap.
|
||||
Under heavy xdist load (40 parallel workers, 6-shard CI), the full
|
||||
sequence can exceed 10s. Default timeout is generous to avoid CI
|
||||
flakes; in practice the wait returns in <1s on quiet hosts.
|
||||
"""
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
if not _pgid_still_alive(pgid):
|
||||
@@ -166,9 +172,11 @@ def test_wait_for_process_kills_subprocess_on_keyboardinterrupt():
|
||||
assert ret == 1, f"SetAsyncExc returned {ret}, expected 1"
|
||||
|
||||
# Give the worker a moment to: hit the exception at the next poll,
|
||||
# run the except-block cleanup (_kill_process), and exit.
|
||||
t.join(timeout=5.0)
|
||||
assert not t.is_alive(), "worker didn't exit within 5 s of the interrupt"
|
||||
# run the except-block cleanup (_kill_process), and exit. Under
|
||||
# xdist load the SIGTERM → 3s wait → SIGKILL chain can take longer
|
||||
# than 5s before the worker's join() returns; bumped to 15s.
|
||||
t.join(timeout=15.0)
|
||||
assert not t.is_alive(), "worker didn't exit within 15 s of the interrupt"
|
||||
|
||||
# The critical assertion: the subprocess GROUP must be dead. Not
|
||||
# just the bash wrapper — the 'sleep 30' child too. Under xdist load,
|
||||
|
||||
Reference in New Issue
Block a user