When the DRM device is lost transiently, for example when the amdgpu driver
performs a GPU reset and recovers ("device wedged, but recovered through
reset"), acquiring a renderer in `SurfaceThreadState::redraw` fails. The three
`renderer`/`single_renderer` calls used `.unwrap()`, so the failure panicked
the surface's render thread. Because that panic unwinds across the EGL/GBM FFI
boundary it aborts the whole process (SIGABRT), dropping the session back to
the greeter and closing every running application.
`redraw` is already fallible and its caller logs the error and reschedules a
redraw, and the frame-submission path (`render_frame`/`queue_frame`) already
propagates errors the same way. Propagate the renderer-acquisition errors too,
so a transient device loss is retried on the next frame, once the device is
back, instead of taking down the compositor.
Relates to #649.
Assisted by an AI coding tool; I have reviewed and fully understand the change
and am able to maintain it and respond to review.
display_configuration() issued an ALLOW_MODESET atomic commit on every udev
event, even when no planes needed detaching. The empty commit is a no-op
modeset that interferes with in-flight commits on the render thread, and it
fails with EPERM whenever we don't hold DRM master (e.g. a render-only
secondary GPU), aborting device enumeration.
Track whether any property was added and skip the commit when there is nothing
to do. A commit that does have changes still propagates its error as before, so
a genuine failure to apply cleanup is not silently swallowed.
https://github.com/pop-os/cosmic-comp/issues/2375
Authored by Claude Opus 4.8
Reviewed by Timo Strunk <Timo.Strunk@gmail.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
When smithay's DrmDevice::activate() returns an error (e.g. because
drmSetMaster() failed during VT resume), resume_session() was only
logging the error and then proceeding to resume DRM leases on a device
that has no master. This leaves the compositor in a broken render state
where every atomic page flip returns EPERM.
Fix: continue to the next device on activate() failure. The device
surface stays inactive and lease resume is deferred. On the next
ActivateSession event (next VT switch to this TTY), resume_session()
will be called again and activate() will succeed once the seat manager
has granted DRM master.
Reproducer: hybrid GPU system (Intel iGPU render + NVIDIA display),
VT switch away from and back to COSMIC session. On return, drmSetMaster
races the seat notification; activate() fails; without this fix
the compositor floods journald with EPERM at 60fps.
Fixes: pop-os#2331, pop-os#2302
Fixes https://github.com/pop-os/cosmic-epoch/issues/2978.
This reverses the part of
ca00df0b37 that made it
only try import on the advertised GPU. But this version avoids
initializing an EGL context simply to re-check the supported texture
formats.
The logic `age_for_buffer` used seems to be a misinterpretation of the
protocol.
The wording is a little unclear, but it seems tracking buffer age is the
responsibility of the client, and the client is required to accumulate
damage and pass it in `damage_buffer`.
Our clients initially weren't doing that correctly. I updated
xdg-desktop-portal-cosmic to use `damage_buffer` after testing on
wlroots, and cosmic-workspaces was recently updated as well.
The important change here is that we now apply the additional damage
first, instead of using `.extend()` to add it after other elements. This
is important since `OutputDamageTracker` will ignore our damage elements
if there are behind an element with an opaque region.
This also makes things a bit simpler, especially `take_screencopy_frames()`,
which no longer needs a mutable references to extend then truncate.
The implementation of `OutputDamageTracker` isn't entirely clear, but as
far as I can tell this is intended to work, and it seems to work in some
testing.
This doesn't change much, since the Smithay implementation is based on
the `cosmic-comp` version, but made more generic. We provide our own
implementation for our workspace capture protocol, but otherwise Smithay
handles the boilerplate now.
This should not cause any change in behavior.
Previously we ignored when we had no output configuration
**and** failed to apply the automatically created one.
This leads to two problems:
- If this happens on startup, we end up with no outputs being added to the shell and we quit.
- If this happens later, we might end up in an inconsistent state, where the shell thinks we have an output, when it didn't light up for similar reasons.
Thus `read_outputs` is failable and handling that very much depends on
the where is was called from, because `read_outputs` doesn't know what
configuration was active before.
Thus make it failable and provide useful mitigations everywhere
possible:
- Try to enable just one output in case we fail on startup.
- Don't enable any additional outputs, when we fail on hotplug.
- Log the error like previously in any other case (and come up with more
mitigations, once we understand these cases better).