Compare commits

...
Author SHA1 Message Date
nmzik 83152c8cad guest_gpu: remove memory-unmap submission deadlock + remove legacy agc buffering 2026-08-02 13:17:53 +02:00
nmzik 35a89d616a implement dynamic 2026-08-02 10:50:28 +02:00
nmzik 59b8fad341 graphics: support 3D color render targets 2026-08-02 09:38:57 +02:00
nmzik b877b4be9c graphics: perf - batch GPU page watcher updates at 4 MiB granularity 2026-08-02 08:39:50 +02:00
nmzik 9da7fc5dd6 renderer: minor optimizations 2026-08-02 08:34:44 +02:00
nmzik da0d33224d renderer: eliminate extra copy for small streaming buffers 2026-08-02 08:34:44 +02:00
nmzik f831e60412 agc: new abis 2026-08-02 07:59:23 +02:00
nmzik 84236d1f87 agc: new abi 2026-08-02 07:59:22 +02:00
Stefanos Costaandnmzik 302b579779 loader: zero unresolved scalar floating-point returns
Extracted from 3db2b3c5c5e1a26a861df7ebcacd9ccb8c484420 in KytyPS5/KytyPS5#147.
2026-08-02 06:35:53 +02:00
Stefanos Costaandnmzik 0b6bf01b36 kernel: preserve microsecond wall-clock resolution
Extracted from 3db2b3c5c5e1a26a861df7ebcacd9ccb8c484420 in KytyPS5/KytyPS5#147.
2026-08-02 06:35:52 +02:00
Stefanos Costaandnmzik 66f640527d audio: fix pacing and AudioOut2 port lifetime
Extracted from 6a60f1b17481a0e5e14242c0fb4dc22f963545e1 in KytyPS5/KytyPS5#147.
2026-08-02 06:35:52 +02:00
nmzik 4631b96178 perf(gpu): run dirty-page validation only in debug builds 2026-08-02 04:58:56 +02:00
43f64e4ab4 Register remaining regression tests with CTest (#24)
Register regression tests with CTest

Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-08-02 04:54:04 +02:00
IdyllizeandGitHub e63f5b7d5c cmake: preserve spaces in clang-cl linker paths (#26)
Pass linker flags as individual options so CMake keeps the PDB and lld map paths intact when the build directory contains spaces.
2026-08-02 04:45:28 +02:00
nikosszzzandnmzik fa7c3c01bf fix: guard Linux memory fixes to only Linux 2026-08-02 03:43:56 +02:00
nikosszzzandnmzik 89651f6f59 kernel/memory: reserve only available guest address ranges on Linux
Reserve only free guest address ranges
2026-08-02 03:43:56 +02:00
nmzik 44d7f2a3e8 shader cfg: handle shared early exits
Duplicate small shared exit tails so each selection gets its own merge block. This keeps overlapping early-exit ladders on structured SPIR-V and adds a regression test.
2026-08-02 03:16:08 +02:00
nmzik 2dcb90066c shader cfg: normalize loop structure
Give loops one header and one continue path before SPIR-V generation. This handles conditional headers and multiple latches without falling back to a dispatcher.
2026-08-02 03:15:26 +02:00
nmzik 51a33cc363 shader cfg: handle loop control branches
Keep simple break, continue, and repeat branches in structured control flow. Split conflicting merge blocks and add regression tests for nested loop exits.
2026-08-02 03:14:39 +02:00
nikosszzzandnmzik ed84370786 fix(libc): run thread-local destructors
Why: Thread-atexit registrations were discarded, leaving objects alive after their guest TLS storage was released.

What: Store registrations per host thread and run them in LIFO order before pthread keys and guest TLS are destroyed.

Why safe: Only callbacks registered on the exiting thread run, once, before existing teardown continues.
2026-08-02 02:48:56 +02:00
Claxtenandnmzik 43f30d3ab2 graphics: shader: ignore unused sampler border state
* Sampler dword 3 only matters when a clamp mode uses border color
  (values >= 4). When no border mode is active, dword 3 is unused
  but can still vary across loop iterations due to wave-lane spills.
  This makes resource tracking think the descriptor is dynamic and
  fail with "unsupported GPU selection".

* Fix by zeroing dword 3 when all clamp modes are non-border.

Signed-off-by: Claxten <claxten10@gmail.com>
2026-08-02 02:42:15 +02:00
Stepz97andGitHub 0838142abd macOS: anchor the guest address space in full-emulator test targets (#143)
fix(cmake): anchor the macOS guest address space for all full-emulator tests

Every target created by add_kyty_full_emulator_test links against the
full kyty_emulator sources, so it drags in the same 620 GiB .zerofill
guest address space segments as the emulator itself. Only the emulator
target and virtual_memory_allocation_tests had the linker flags that
anchor those segments; every other full-emulator test target got the
segments without the anchoring, and the kernel killed them on exec
(posix_spawn EIO / SIGKILL) before main() ever ran.

Move the configure_macos_guest_address_space() call into
add_kyty_full_emulator_test() itself so every target it creates gets
it automatically, and drop the now-redundant explicit call on
virtual_memory_allocation_tests.
2026-08-02 02:27:32 +02:00
0f550d1fd0 fix: keep hint-less guest mappings at the canonical PS5 base (fixes the #135 macOS regression) (#138)
* fix: keep hint-less guest mappings at the canonical PS5 base

FindGuestFreeRange searched the low system-managed range first for
mappings with no address hint, so the first hint-less direct-memory map
could land as low as 0x200000. The PS5 kernel never places hint-less
user mappings below 0x200000000 and guest code relies on that: Sony's
libc maps 4 MiB of direct memory for its internal heap, fails its
mspace setup when the returned address is that low, and the first
malloc then dereferences a null mspace (a read at 0x38, the mspace
magic check). On macOS this made Raiden III crash on the main guest
thread a couple of seconds after boot, 100 percent reproducible with
--printf-direction Silent.

Search from the canonical base first, fall back to the user range, and
keep the low system-managed range only as a last resort. The mmap path
already anchored hint-less searches at 0x200000000; this aligns the
shared search helper with it.

Adds two regression tests: the libc-shaped allocation must come back at
or above the canonical base and hold writes, and direct-memory content
must survive an unmap and remap of the same physical range.

* macos: make the fatal-report memory dumps fault-safe

IsReadableRange returned true for any nonzero address on macOS, so the
fatal report's guest memory dumps dereferenced whatever the crashed
thread had in its registers. A fault inside the reporter re-enters the
signal handler and wedges the reporting thread, which hid real guest
crashes whenever logging was enabled: the game kept running with a dead
thread and the report was never completed.

Walk the Mach regions covering the range and require read permission
before dumping, the same contract the Linux implementation provides.

* do not fallthrough HOST_SYSTEM_MANAGED_MIN

---------

Co-authored-by: nmzik <Nmzik@mail.ru>
2026-08-02 02:23:12 +02:00
Claxtenandnmzik c1a5927036 graphics: pm4: accept trailing PM4 type-2 packets
* A one-dword type-2 NOP is a valid packet tail. Parse it normally instead of aborting command-buffer dumps.

Signed-off-by: Claxten <claxten10@gmail.com>
2026-08-02 02:07:37 +02:00
Claxtenandnmzik bc436548a9 graphics: support packed 10-10-10-2 uint buffers
Signed-off-by: Claxten <claxten10@gmail.com>
2026-08-02 01:44:21 +02:00
nmzik a65d17a5d6 renderer: skip debug checks when stencil&depth is not active 2026-08-01 00:22:24 +02:00
nmzik c690aeea62 add HiS PM4 handler, accept Ngs2CustomMastering 2026-08-01 00:12:42 +02:00
nmzik f830d6b2e4 renderer: remove legacy code left from refactoring 2026-07-31 23:13:48 +02:00
nmzik d68a477276 renderer: broaden compatibility 2026-07-31 22:24:56 +02:00
nmzik 8977d4d2f0 shader: add descriptor log 2026-07-31 21:29:26 +02:00
nmzik 4b4e3bf3cf pm4: implement missing selectors 2026-07-31 19:45:47 +02:00
nmzik a0bb129f02 SaveData: stop escaping root directory 2026-07-31 18:53:54 +02:00
nmzik 846002c5eb shader: fix invalid texture descriptors and add missing format 2026-07-31 17:30:31 +02:00
nmzik d8a4c83cc7 format src and tests with clang-format 2026-07-31 11:36:12 +02:00
nmzik 68be13345a fix(shader): support multisampled depth image loads 2026-07-31 11:36:12 +02:00
nmzik 167da0abe0 shader: fix readlane/writelane for inactive host lanes 2026-07-31 11:36:12 +02:00
nmzik 48c31d61ee Implement VideoDec2 2026-07-31 11:36:12 +02:00
nmzik 212282d693 fix(shader): preserve packed UINT16 MRT exports 2026-07-31 11:36:12 +02:00
nmzik 6bca35d1f5 renderer: broaden compatibility 2026-07-31 11:36:12 +02:00
M. AbdullahandGitHub e4ad5fc988 docs: add macOS build and run instructions (#137)
The README had macOS badges and an experimental-support note but no build,
run, or system-requirement information for the platform. Document the
Rosetta 2 / MoltenVK setup, the x86-64 configure invocation, the Qt
universal-build requirement, MoltenVK installation and signing, and the
SDL_VULKAN_LIBRARY variable needed at run time.
2026-07-31 05:32:29 +02:00
nmzik c0d3d261ea add TextToSpeech2 stubs 2026-07-31 04:05:50 +02:00
3b75a5659a shader: specialize cube image descriptors (#134)
* shader: specialize cube image descriptors

Track whether image descriptors refer to cube maps during resource specialization, and apply the coordinate offset conversion when sampling cube maps as 2D image arrays in SPIR-V emission.

* shader: fix cube array coordinate lowering

---------

Co-authored-by: nmzik <Nmzik@mail.ru>
2026-07-31 03:59:10 +02:00
M. AbdullahandGitHub d475387171 macOS: enable guest signal dispatch on the target thread (#136)
macos: enable guest signal dispatch on the target thread

The POSIX signal-dispatch path (pthread_kill based, added with the Linux
port) was compiled out on macOS, leaving KernelRaiseException to run the
guest handler on the calling thread. IL2CPP's garbage collector raises its
stop-the-world signal at every managed thread and each handler parks its
own thread until resume, so the collector parked itself and every Unity
title froze on the first collection.

Enable the same delivery path on macOS:
- translate between the Darwin mcontext (uc_mcontext->__ss) and the guest
  ucontext in CreateSignalUcontextFromHost/ApplySignalUcontextToHost
- use SIGUSR1 as the host dispatch signal (macOS has no realtime signals)
- block the dispatch signal inside the host fault handler so a suspend
  request cannot preempt fault resolution between the protection fix and
  the retry

Windows and Linux are unchanged.
2026-07-31 03:47:04 +02:00
nmzikandGitHub 2f5396c6a5 Rework guest memory tracking/virtual address space/direct and flexible memory (#135)
* Rework guest memory tracking

* add unknwon flag

* Fix macOS guest address-space reservation
2026-07-31 03:07:17 +02:00
156 changed files with 27371 additions and 26202 deletions
+17 -3
View File
@@ -83,7 +83,12 @@ jobs:
- name: Build
shell: cmd
run: |
cmake --build _Build/windows --target launcher --parallel
cmake --build _Build/windows --target launcher audio_out2_port_tests virtual_memory_allocation_tests --parallel
- name: Test
shell: cmd
run: |
ctest --test-dir _Build/windows --output-on-failure -R "^(audio_out2_port|virtual_memory_allocation)$"
- name: Install
shell: cmd
@@ -153,7 +158,15 @@ jobs:
- name: Build
shell: bash
run: |
cmake --build _Build/macos --target launcher --parallel
cmake --build _Build/macos \
--target launcher audio_out2_port_tests virtual_memory_allocation_tests \
--parallel
- name: Test
shell: bash
run: |
ctest --test-dir _Build/macos --output-on-failure \
-R '^(audio_out2_port|virtual_memory_allocation)$'
- name: Install
shell: bash
@@ -284,13 +297,14 @@ jobs:
run: |
cmake --build _Build/linux \
--target launcher page_manager_tests memory_tracker_tests \
audio_out2_port_tests virtual_memory_allocation_tests \
--parallel
- name: Test
shell: bash
run: |
ctest --test-dir _Build/linux --output-on-failure \
-R '^(page_manager|memory_tracker)$'
-R '^(audio_out2_port|page_manager|memory_tracker|virtual_memory_allocation)$'
- name: Install
shell: bash
+65 -5
View File
@@ -27,8 +27,9 @@ Development is focused on compatibility and boot reliability.
Windows is the primary platform and receives the most testing. Linux builds and runs; see
[Building on Linux](#building-on-linux).
macOS support is experimental. Compatibility with the same games on Windows and macOS has not yet
been tested.
macOS support is experimental. The emulator is built for x86-64 and runs on Apple Silicon under
Rosetta 2, with Vulkan provided by MoltenVK. A small number of titles have been verified in-game
on Apple Silicon hardware; see [Building on macOS](#building-on-macos).
## Bugs and Issues
@@ -114,9 +115,10 @@ the Vulkan/SPIR-V validation rules.
### System requirements
- Windows 10 version 1803, or a current Linux distribution
- A 64-bit x86 processor
- A Vulkan 1.3-capable GPU with current drivers
- Windows 10 version 1803, a current Linux distribution, or macOS on Apple Silicon
- A 64-bit x86 processor (on macOS, an Apple Silicon processor with Rosetta 2)
- A Vulkan 1.3-capable GPU with current drivers (on macOS, Vulkan is provided by the bundled
MoltenVK)
### Build requirements (Windows)
@@ -188,6 +190,56 @@ time.
Note that the CMake source root is `src`, not the repository root.
### Building on macOS
macOS builds target x86-64 and run under Rosetta 2 on Apple Silicon, so the PS5's x86-64 game
code executes through the same translation layer as the emulator itself. Prebuilt archives are
attached to releases; the steps below are for building from source.
Requirements:
- An Apple Silicon Mac with Rosetta 2 installed (`softwareupdate --install-rosetta`)
- Xcode (or the Command Line Tools)
- Homebrew packages: `brew install cmake ninja glslang`
- Qt 6 (Concurrent, Network, Widgets) with x86-64 support. The official Qt installation is
universal and works; Homebrew's Qt is arm64-only and will not link
```bash
git submodule update --init --recursive
cmake -S src -B _Build/macos -G Ninja -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_OSX_ARCHITECTURES=x86_64 \
-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++ \
-DCMAKE_PREFIX_PATH="$Qt6_DIR"
cmake --build _Build/macos --target launcher --parallel
cmake --install _Build/macos --prefix _Build/macos/install
```
The build re-signs `kyty_emulator` with the JIT entitlements it needs to execute translated
guest code; no manual signing step is required.
Vulkan comes from MoltenVK. Download `MoltenVK-macos.tar` from the
[MoltenVK releases](https://github.com/KhronosGroup/MoltenVK/releases), then copy
`MoltenVK/dynamic/dylib/macOS/libMoltenVK.dylib` next to `kyty_emulator` and ad-hoc sign it:
```bash
codesign --force --sign - _Build/macos/install/libMoltenVK.dylib
```
Release archives already include a signed `libMoltenVK.dylib`.
### Regression tests
Build every regression executable and run the registered tests with:
```powershell
cmake --build _Build/windows --target kyty_tests
ctest --test-dir _Build/windows --output-on-failure
```
Use `_Build/linux` instead of `_Build/windows` for a Linux build.
### Visual Studio Code
A ready-made Visual Studio Code setup is included in [`.vscode`](.vscode). It configures CMake
@@ -233,6 +285,14 @@ The emulator can also be started directly with a legally obtained game directory
./_Build/linux/install/kyty_emulator --game "/games/ExampleGame"
```
On macOS, point SDL at the MoltenVK library explicitly; the hardened runtime prevents it from
being picked up from the executable's directory:
```bash
cd _Build/macos/install
SDL_VULKAN_LIBRARY="$PWD/libMoltenVK.dylib" ./kyty_emulator --game "/games/ExampleGame"
```
Run `kyty_emulator --help` to see the available graphics, logging, validation, profiling, and
debugging options.
+59 -2
View File
@@ -312,6 +312,19 @@ function(add_kyty_full_emulator_test target source)
target_link_libraries(${target} onecore)
add_custom_command(TARGET ${target} POST_BUILD COMMAND ${CMAKE_COMMAND} -E copy_if_different "${KYTY_THIRD_PARTY_DIR}/winpthread/bin/libwinpthread-1.dll" $<TARGET_FILE_DIR:${target}>/libwinpthread-1.dll)
endif()
# The macOS x86_64 guest address space needs its .zerofill segments anchored
# by linker flags, or the kernel kills the binary on load (posix_spawn EIO).
configure_macos_guest_address_space(${target})
endfunction()
function(configure_macos_guest_address_space target)
if(APPLE AND (CMAKE_OSX_ARCHITECTURES STREQUAL "x86_64" OR
(NOT CMAKE_OSX_ARCHITECTURES AND CMAKE_SYSTEM_PROCESSOR MATCHES "^(x86_64|AMD64)$")))
target_sources(${target} PRIVATE kernel/macosGuestAddressSpace.cpp)
target_compile_definitions(${target} PRIVATE KYTY_LINKED_GUEST_ADDRESS_SPACE=1)
target_link_options(${target} PRIVATE
-Wl,-ld_classic,-no_pie,-no_fixup_chains,-no_huge,-pagezero_size,0x40000,-segaddr,SYSTEM_MANAGED,0x40000,-segaddr,SYSTEM_RESERVED,0x7ffffc000,-segaddr,USER_AREA,0x7000000000,-image_base,0x700000000000)
endif()
endfunction()
add_kyty_full_emulator_test(shader_cfg_tests ../tests/shaderCfgTests.cpp)
@@ -319,6 +332,7 @@ add_kyty_full_emulator_test(shader_cfg_tests ../tests/shaderCfgTests.cpp)
add_executable(scalar_provenance_tests EXCLUDE_FROM_ALL
../tests/ScalarProvenanceTests.cpp
graphics/host_gpu/hostMemory.cpp
graphics/shader/recompiler/ir/ReadLaneElimination.cpp
graphics/shader/recompiler/ir/ScalarProvenance.cpp
graphics/shader/recompiler/ir/SrtWalker.cpp
)
@@ -331,6 +345,11 @@ add_executable(page_manager_tests EXCLUDE_FROM_ALL
)
target_include_directories(page_manager_tests PRIVATE ${inc_headers})
add_executable(bit_array_tests EXCLUDE_FROM_ALL
../tests/BitArrayTests.cpp
)
target_include_directories(bit_array_tests PRIVATE ${inc_headers})
add_executable(memory_tracker_tests EXCLUDE_FROM_ALL
../tests/MemoryTrackerTests.cpp
graphics/host_gpu/pageManager.cpp
@@ -338,7 +357,6 @@ add_executable(memory_tracker_tests EXCLUDE_FROM_ALL
)
target_link_libraries(memory_tracker_tests fmt::fmt common)
target_include_directories(memory_tracker_tests PRIVATE ${inc_headers})
target_compile_definitions(memory_tracker_tests PRIVATE KYTY_MEMORY_TRACKER_TESTS=1)
add_executable(shader_vertex_metadata_tests EXCLUDE_FROM_ALL
../tests/ShaderVertexMetadataTests.cpp
@@ -381,6 +399,14 @@ add_executable(resource_mutex_tests EXCLUDE_FROM_ALL
target_link_libraries(resource_mutex_tests common)
target_include_directories(resource_mutex_tests PRIVATE ${inc_headers})
add_executable(audio_out2_port_tests EXCLUDE_FROM_ALL
../tests/AudioOut2PortTests.cpp
libs/libAudio2.cpp
loader/timer.cpp
)
target_link_libraries(audio_out2_port_tests common fmt::fmt)
target_include_directories(audio_out2_port_tests PRIVATE ${inc_headers})
add_executable(event_queue_lifetime_tests EXCLUDE_FROM_ALL
../tests/EventQueueLifetimeTests.cpp
kernel/eventQueue.cpp
@@ -431,12 +457,21 @@ if(NOT KYTY_CLANG_CL)
endif()
if(BUILD_TESTING)
add_test(NAME shader_cfg COMMAND $<TARGET_FILE:shader_cfg_tests>)
add_test(NAME scalar_provenance COMMAND $<TARGET_FILE:scalar_provenance_tests>)
add_test(NAME image_page_table COMMAND $<TARGET_FILE:image_page_table_tests>)
add_test(NAME memory_tracker COMMAND $<TARGET_FILE:memory_tracker_tests>)
add_test(NAME page_manager COMMAND $<TARGET_FILE:page_manager_tests>)
add_test(NAME bit_array COMMAND $<TARGET_FILE:bit_array_tests>)
add_test(NAME shader_vertex_metadata COMMAND $<TARGET_FILE:shader_vertex_metadata_tests>)
add_test(NAME shader_stage_runtime COMMAND $<TARGET_FILE:shader_stage_runtime_tests>)
add_test(NAME resource_tracking COMMAND $<TARGET_FILE:resource_tracking_tests>)
add_test(NAME resource_mutex COMMAND $<TARGET_FILE:resource_mutex_tests>)
add_test(NAME event_queue_lifetime COMMAND $<TARGET_FILE:event_queue_lifetime_tests>)
add_test(NAME audio_out2_port COMMAND $<TARGET_FILE:audio_out2_port_tests>)
add_test(NAME shader_recompiler_compute COMMAND $<TARGET_FILE:shader_recompiler_compute_tests>)
add_test(NAME virtual_memory_allocation
COMMAND $<TARGET_FILE:virtual_memory_allocation_tests>)
add_test(NAME command_scheduler_timeline
COMMAND $<TARGET_FILE:shader_recompiler_compute_tests> --scheduler-only)
add_test(NAME stream_buffer_ring
@@ -466,10 +501,27 @@ if(BUILD_TESTING)
add_test(NAME buffer_cache_ranges
COMMAND $<TARGET_FILE:shader_recompiler_compute_tests> --buffer-cache-range-only)
endif()
add_custom_target(kyty_tests DEPENDS
shader_cfg_tests
scalar_provenance_tests
image_page_table_tests
memory_tracker_tests
page_manager_tests
bit_array_tests
shader_vertex_metadata_tests
shader_stage_runtime_tests
resource_tracking_tests
resource_mutex_tests
event_queue_lifetime_tests
shader_recompiler_compute_tests
virtual_memory_allocation_tests
)
endif()
add_executable(kyty_emulator main.cpp ${kyty_emulator_src})
configure_macos_guest_address_space(kyty_emulator)
target_link_libraries(kyty_emulator ${kyty_emulator_link_libraries})
if (WIN32)
@@ -498,7 +550,12 @@ set(KYTY_EMULATOR_MAP_LINK_PATH "${CMAKE_CURRENT_BINARY_DIR}/${KYTY_EMULATOR_MAP
set(KYTY_EMULATOR_PDB_LINK_PATH "${CMAKE_CURRENT_BINARY_DIR}/kyty_emulator.pdb")
if(KYTY_CLANG_CL)
set_target_properties(kyty_emulator PROPERTIES LINK_FLAGS "/DYNAMICBASE:NO /DEBUG:FULL /PDB:${KYTY_EMULATOR_PDB_LINK_PATH} /lldmap:${KYTY_EMULATOR_MAP_LINK_PATH}")
target_link_options(kyty_emulator PRIVATE
"/DYNAMICBASE:NO"
"/DEBUG:FULL"
"/PDB:${KYTY_EMULATOR_PDB_LINK_PATH}"
"/lldmap:${KYTY_EMULATOR_MAP_LINK_PATH}"
)
add_custom_command(TARGET kyty_emulator POST_BUILD COMMAND ${CMAKE_COMMAND} -E copy_if_different "${KYTY_THIRD_PARTY_DIR}/winpthread/bin/libwinpthread-1.dll" $<TARGET_FILE_DIR:kyty_emulator>/libwinpthread-1.dll)
elseif(WIN32 OR LINUX)
set_target_properties(kyty_emulator PROPERTIES LINK_FLAGS "${KYTY_LD_OPTIONS} -Wl,-Map=${KYTY_EMULATOR_MAP_LINK_PATH}")
+260
View File
@@ -0,0 +1,260 @@
#ifndef EMULATOR_SRC_COMMON_BITARRAY_H_
#define EMULATOR_SRC_COMMON_BITARRAY_H_
#include <array>
#include <bit>
#include <cstddef>
#include <cstdint>
#include <iterator>
#include <utility>
namespace Common {
template <size_t N>
class BitArray final {
static_assert(N != 0, "BitArray size must be nonzero");
static_assert(N % 64 == 0, "BitArray size must be a multiple of 64 bits");
static constexpr size_t BITS_PER_WORD = 64;
static constexpr size_t WORD_COUNT = N / BITS_PER_WORD;
public:
using Range = std::pair<size_t, size_t>;
class Iterator final {
public:
using iterator_category = std::forward_iterator_tag;
using value_type = Range;
using difference_type = std::ptrdiff_t;
using pointer = const Range*;
using reference = const Range&;
Iterator(const BitArray& bits, size_t start)
: m_bits(bits), m_range(bits.FirstRangeFrom(start)) {}
Iterator& operator++() {
m_range = m_bits.FirstRangeFrom(m_range.second);
return *this;
}
[[nodiscard]] bool operator==(const Iterator& other) const {
return &m_bits == &other.m_bits && m_range == other.m_range;
}
[[nodiscard]] bool operator!=(const Iterator& other) const { return !(*this == other); }
[[nodiscard]] reference operator*() const { return m_range; }
[[nodiscard]] pointer operator->() const { return &m_range; }
private:
const BitArray& m_bits;
Range m_range;
};
using const_iterator = Iterator;
constexpr BitArray() = default;
constexpr BitArray(const BitArray& other, size_t start, size_t end) {
if (start >= end || end > N) {
return;
}
const auto first_word = start / BITS_PER_WORD;
const auto last_word = (end - 1) / BITS_PER_WORD;
const auto start_bit = start % BITS_PER_WORD;
const auto end_bit = (end - 1) % BITS_PER_WORD;
const auto start_mask = ~uint64_t {0} << start_bit;
const auto end_mask =
end_bit == BITS_PER_WORD - 1 ? ~uint64_t {0} : (uint64_t {1} << (end_bit + 1)) - 1;
if (first_word == last_word) {
m_data[first_word] = other.m_data[first_word] & start_mask & end_mask;
return;
}
m_data[first_word] = other.m_data[first_word] & start_mask;
for (auto word = first_word + 1; word < last_word; word++) {
m_data[word] = other.m_data[word];
}
m_data[last_word] = other.m_data[last_word] & end_mask;
}
[[nodiscard]] constexpr bool Get(size_t index) const {
return (m_data[index / BITS_PER_WORD] & (uint64_t {1} << (index % BITS_PER_WORD))) != 0;
}
constexpr void Set(size_t index) {
m_data[index / BITS_PER_WORD] |= uint64_t {1} << (index % BITS_PER_WORD);
}
constexpr void Unset(size_t index) {
m_data[index / BITS_PER_WORD] &= ~(uint64_t {1} << (index % BITS_PER_WORD));
}
constexpr void SetRange(size_t start, size_t end) {
if (start >= end || end > N) {
return;
}
const auto first_word = start / BITS_PER_WORD;
const auto last_word = (end - 1) / BITS_PER_WORD;
const auto start_bit = start % BITS_PER_WORD;
const auto end_bit = (end - 1) % BITS_PER_WORD;
const auto start_mask = ~uint64_t {0} << start_bit;
const auto end_mask =
end_bit == BITS_PER_WORD - 1 ? ~uint64_t {0} : (uint64_t {1} << (end_bit + 1)) - 1;
if (first_word == last_word) {
m_data[first_word] |= start_mask & end_mask;
return;
}
m_data[first_word] |= start_mask;
for (auto word = first_word + 1; word < last_word; word++) {
m_data[word] = ~uint64_t {0};
}
m_data[last_word] |= end_mask;
}
constexpr void UnsetRange(size_t start, size_t end) {
if (start >= end || end > N) {
return;
}
const auto first_word = start / BITS_PER_WORD;
const auto last_word = (end - 1) / BITS_PER_WORD;
const auto start_bit = start % BITS_PER_WORD;
const auto end_bit = (end - 1) % BITS_PER_WORD;
const auto start_mask = (uint64_t {1} << start_bit) - 1;
const auto end_mask =
end_bit == BITS_PER_WORD - 1 ? uint64_t {0} : ~((uint64_t {1} << (end_bit + 1)) - 1);
if (first_word == last_word) {
m_data[first_word] &= start_mask | end_mask;
return;
}
m_data[first_word] &= start_mask;
for (auto word = first_word + 1; word < last_word; word++) {
m_data[word] = 0;
}
m_data[last_word] &= end_mask;
}
constexpr void Clear() { m_data.fill(0); }
constexpr void Fill() { m_data.fill(~uint64_t {0}); }
[[nodiscard]] constexpr bool None() const {
uint64_t combined = 0;
for (const auto word: m_data) {
combined |= word;
}
return combined == 0;
}
[[nodiscard]] constexpr bool Any() const { return !None(); }
[[nodiscard]] constexpr Range FirstRangeFrom(size_t start) const {
if (start >= N) {
return {N, N};
}
auto word_index = start / BITS_PER_WORD;
auto word = m_data[word_index] & (~uint64_t {0} << (start % BITS_PER_WORD));
while (word == 0) {
word_index++;
if (word_index == WORD_COUNT) {
return {N, N};
}
word = m_data[word_index];
}
const auto first = word_index * BITS_PER_WORD + std::countr_zero(word);
const auto first_bit = first % BITS_PER_WORD;
const auto first_ones =
static_cast<size_t>(std::countr_one(m_data[word_index] >> first_bit));
if (first_bit + first_ones < BITS_PER_WORD) {
return {first, first + first_ones};
}
for (word_index++; word_index < WORD_COUNT; word_index++) {
word = m_data[word_index];
if (word != ~uint64_t {0}) {
return {first, word_index * BITS_PER_WORD + std::countr_one(word)};
}
}
return {first, N};
}
[[nodiscard]] constexpr Range FirstRange() const { return FirstRangeFrom(0); }
[[nodiscard]] constexpr Range LastRangeFrom(size_t end) const {
if (end == 0) {
return {0, 0};
}
if (end > N) {
end = N;
}
auto word_index = (end - 1) / BITS_PER_WORD;
const auto end_bit = (end - 1) % BITS_PER_WORD;
const auto end_mask =
end_bit == BITS_PER_WORD - 1 ? ~uint64_t {0} : (uint64_t {1} << (end_bit + 1)) - 1;
auto word = m_data[word_index] & end_mask;
while (word == 0) {
if (word_index == 0) {
return {0, 0};
}
word = m_data[--word_index];
}
const auto empty_bits = static_cast<size_t>(std::countl_zero(word));
const auto ones = static_cast<size_t>(std::countl_one(word << empty_bits));
const auto last = (word_index + 1) * BITS_PER_WORD - empty_bits;
if (empty_bits + ones < BITS_PER_WORD) {
return {last - ones, last};
}
while (word_index != 0) {
word = m_data[--word_index];
if (word != ~uint64_t {0}) {
return {(word_index + 1) * BITS_PER_WORD - std::countl_one(word), last};
}
}
return {0, last};
}
[[nodiscard]] constexpr Range LastRange() const { return LastRangeFrom(N); }
[[nodiscard]] const_iterator begin() const { return Iterator(*this, 0); }
[[nodiscard]] const_iterator end() const { return Iterator(*this, N); }
constexpr BitArray& operator^=(const BitArray& other) {
for (size_t word = 0; word < WORD_COUNT; word++) {
m_data[word] ^= other.m_data[word];
}
return *this;
}
[[nodiscard]] constexpr BitArray operator^(const BitArray& other) const {
auto result = *this;
result ^= other;
return result;
}
[[nodiscard]] constexpr BitArray operator~() const {
auto result = *this;
for (auto& word: result.m_data) {
word = ~word;
}
return result;
}
private:
std::array<uint64_t, WORD_COUNT> m_data {};
};
} // namespace Common
#endif // EMULATOR_SRC_COMMON_BITARRAY_H_
+10 -6
View File
@@ -175,9 +175,9 @@ static void SignalHandler(int sig, siginfo_t* si, void* uctx) {
}
g_in_exception_filter = true;
auto* uc = static_cast<ucontext_t*>(uctx);
const auto* mc = uc->uc_mcontext;
const auto& ss = mc->__ss;
auto* uc = static_cast<ucontext_t*>(uctx);
const auto* mc = uc->uc_mcontext;
const auto& ss = mc->__ss;
ExceptionInfo info {};
info.exception_address = ss.__rip;
@@ -214,7 +214,7 @@ static void SignalHandler(int sig, siginfo_t* si, void* uctx) {
FailFast("host exception callback is null");
}
const bool resolved = handler(info);
const bool resolved = handler(info);
g_in_exception_filter = false;
if (resolved) {
@@ -255,8 +255,8 @@ static void SignalHandler(int signal_number, siginfo_t* signal_info, void* nativ
info.native_context = context;
if (signal_number == SIGSEGV || signal_number == SIGBUS) {
info.type = ExceptionType::AccessViolation;
const auto error_code = static_cast<uint64_t>(gregs[REG_ERR]);
info.type = ExceptionType::AccessViolation;
const auto error_code = static_cast<uint64_t>(gregs[REG_ERR]);
if ((error_code & PAGE_FAULT_ERROR_INSTRUCTION) != 0) {
info.access_violation_type = AccessViolationType::Execute;
} else if ((error_code & PAGE_FAULT_ERROR_WRITE) != 0) {
@@ -324,6 +324,10 @@ bool InstallHandler(Handler handler) {
sa.sa_sigaction = SignalHandler;
sa.sa_flags = SA_SIGINFO;
sigemptyset(&sa.sa_mask);
// The guest signal-dispatch path (KernelRaiseException) interrupts threads with
// SIGUSR1; block it while a fault is being resolved so a stop-the-world request
// cannot preempt the handler between the protection fix and the retry.
sigaddset(&sa.sa_mask, SIGUSR1);
// macOS raises SIGBUS for protection faults on some paths and SIGSEGV on others;
// SIGILL covers instructions the host cannot execute (routed to the x64 emulator).
+7 -8
View File
@@ -19,10 +19,10 @@ class LeastRecentlyUsedCache {
public:
[[nodiscard]] size_t Insert(Object object, Tick tick) {
const auto id = Build();
const auto id = Build();
auto& item = m_items[id];
item.object = std::move(object);
item.tick = tick;
item.object = std::move(object);
item.tick = tick;
Attach(item);
return id;
}
@@ -49,8 +49,7 @@ public:
template <typename Function>
void ForEachItemBelow(Tick tick, Function&& function) {
constexpr bool ReturnsBool =
std::is_same_v<std::invoke_result_t<Function, Object>, bool>;
constexpr bool ReturnsBool = std::is_same_v<std::invoke_result_t<Function, Object>, bool>;
for (auto* item = m_first; item != nullptr;) {
if (item->tick > tick) {
return;
@@ -87,10 +86,10 @@ private:
m_last = &item;
return;
}
item.prev = m_last;
item.prev = m_last;
m_last->next = &item;
item.next = nullptr;
m_last = &item;
item.next = nullptr;
m_last = &item;
}
void Detach(Item& item) {
+3 -4
View File
@@ -31,10 +31,9 @@ static bool OnOwnStack() {
if (pthread_getattr_np(pthread_self(), &attr) != 0) {
return false;
}
void* base = nullptr;
size_t size = 0;
const bool ok =
pthread_attr_getstack(&attr, &base, &size) == 0 && base != nullptr && size != 0;
void* base = nullptr;
size_t size = 0;
const bool ok = pthread_attr_getstack(&attr, &base, &size) == 0 && base != nullptr && size != 0;
pthread_attr_destroy(&attr);
if (!ok) {
return false;
+3 -5
View File
@@ -172,8 +172,7 @@ sys_file_t* SysFileCreate(const std::filesystem::path& file_name) {
return ret;
}
sys_file_t* SysFileOpenR(const std::filesystem::path& file_name,
sys_file_cache_type_t cache_type) {
sys_file_t* SysFileOpenR(const std::filesystem::path& file_name, sys_file_cache_type_t cache_type) {
auto* ret = new sys_file_t;
ret->type = SYS_FILE_FILE;
@@ -218,8 +217,7 @@ sys_file_t* SysFileCreate() {
return ret;
}
sys_file_t* SysFileOpenW(const std::filesystem::path& file_name,
sys_file_cache_type_t cache_type) {
sys_file_t* SysFileOpenW(const std::filesystem::path& file_name, sys_file_cache_type_t cache_type) {
auto* ret = new sys_file_t;
auto real_name = get_internal_name(file_name);
@@ -241,7 +239,7 @@ sys_file_t* SysFileOpenW(const std::filesystem::path& file_name,
}
sys_file_t* SysFileOpenRw(const std::filesystem::path& file_name,
sys_file_cache_type_t cache_type) {
sys_file_cache_type_t cache_type) {
auto* ret = new sys_file_t;
auto real_name = get_internal_name(file_name);
+13 -13
View File
@@ -136,8 +136,8 @@ static void* map_anonymous(uintptr_t addr, size_t size, int protect, int flags)
break;
}
const auto hint = (top - step) & ~(LOW_ARENA_GRAIN - 1);
void* ptr = mmap(reinterpret_cast<void*>(hint), size, protect,
flags | MAP_FIXED_NOREPLACE, -1, 0); // NOLINT
void* ptr = mmap(reinterpret_cast<void*>(hint), size, protect, flags | MAP_FIXED_NOREPLACE,
-1, 0); // NOLINT
if (ptr != MAP_FAILED) {
return ptr;
}
@@ -161,8 +161,8 @@ uint64_t SysVirtualAlloc(uint64_t address, uint64_t size, VirtualMemory::Mode mo
if (ptr != MAP_FAILED) {
pthread_mutex_lock(&g_virtual_mutex);
record_alloc(ret_addr, size);
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
for (uintptr_t page = page_start; page <= page_end; page++) {
(*g_protects)[page] = protect;
}
@@ -194,8 +194,8 @@ uint64_t SysVirtualAllocAligned(uint64_t address, uint64_t size, VirtualMemory::
if (ptr != MAP_FAILED && ((ret_addr & (alignment - 1)) != 0)) {
munmap(ptr, size);
ptr = map_anonymous(addr, size + alignment, protect,
MAP_PRIVATE | MAP_ANON | MAP_NORESERVE);
ptr =
map_anonymous(addr, size + alignment, protect, MAP_PRIVATE | MAP_ANON | MAP_NORESERVE);
ret_addr = reinterpret_cast<uintptr_t>(ptr);
if (ptr != MAP_FAILED) {
#if defined(__APPLE__)
@@ -251,8 +251,8 @@ uint64_t SysVirtualAllocAligned(uint64_t address, uint64_t size, VirtualMemory::
pthread_mutex_lock(&g_virtual_mutex);
record_alloc(ret_addr, size);
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
for (uintptr_t page = page_start; page <= page_end; page++) {
(*g_protects)[page] = protect;
}
@@ -266,9 +266,9 @@ uint64_t SysVirtualAllocAligned(uint64_t address, uint64_t size, VirtualMemory::
// the first mapped region at or above `region_addr`; if it begins before the end of the
// requested range, the range overlaps an existing mapping.
static bool is_mapped(void* ptr, size_t length) {
auto query_addr = reinterpret_cast<mach_vm_address_t>(ptr);
mach_vm_address_t region_addr = query_addr;
mach_vm_size_t region_size = 0;
auto query_addr = reinterpret_cast<mach_vm_address_t>(ptr);
mach_vm_address_t region_addr = query_addr;
mach_vm_size_t region_size = 0;
vm_region_basic_info_data_64_t info {};
mach_msg_type_number_t count = VM_REGION_BASIC_INFO_COUNT_64;
mach_port_t object_name = MACH_PORT_NULL;
@@ -337,8 +337,8 @@ bool SysVirtualAllocFixed(uint64_t address, uint64_t size, VirtualMemory::Mode m
if (ptr != MAP_FAILED) {
pthread_mutex_lock(&g_virtual_mutex);
record_alloc(ret_addr, size);
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
for (uintptr_t page = page_start; page <= page_end; page++) {
(*g_protects)[page] = protect;
}
+1 -1
View File
@@ -5,9 +5,9 @@
#include <algorithm>
#include <atomic>
#include <cerrno>
#include <chrono> // IWYU pragma: keep
#include <condition_variable> // IWYU pragma: keep
#include <cerrno>
#include <mutex>
#include <vector>
+2 -4
View File
@@ -11,7 +11,7 @@ template <typename Result, typename... Args>
class UniqueFunction {
class CallableBase {
public:
virtual ~CallableBase() = default;
virtual ~CallableBase() = default;
virtual Result Invoke(Args&&... args) = 0;
};
@@ -20,9 +20,7 @@ class UniqueFunction {
public:
explicit Callable(Function function): m_function(std::move(function)) {}
Result Invoke(Args&&... args) override {
return m_function(std::forward<Args>(args)...);
}
Result Invoke(Args&&... args) override { return m_function(std::forward<Args>(args)...); }
private:
Function m_function;
-19
View File
@@ -58,25 +58,6 @@ bool FlushInstructionCache(uint64_t address, uint64_t size) {
return SysVirtualFlushInstructionCache(address, size);
}
bool PatchReplace(uint64_t vaddr, uint64_t value) {
Mode old_mode {};
Protect(vaddr, 8, Mode::ReadWrite, &old_mode);
auto* ptr = reinterpret_cast<uint64_t*>(vaddr);
bool ret = (*ptr != value);
*ptr = value;
Protect(vaddr, 8, old_mode);
if (IsExecute(old_mode)) {
FlushInstructionCache(vaddr, 8);
}
return ret;
}
} // namespace VirtualMemory
} // namespace Common
-1
View File
@@ -37,7 +37,6 @@ bool Free(uint64_t address);
bool FreeRange(uint64_t address, uint64_t size);
bool Protect(uint64_t address, uint64_t size, Mode mode, Mode* old_mode = nullptr);
bool FlushInstructionCache(uint64_t address, uint64_t size);
bool PatchReplace(uint64_t vaddr, uint64_t value);
} // namespace VirtualMemory
+13 -12
View File
@@ -105,7 +105,7 @@ static void ClearDebugTextureFolder() {
}
}
static void Init(const Config::ConfigOptions& cfg) {
static void Init(const Config::ConfigOptions& cfg, const std::filesystem::path& param_json) {
EXIT_IF(!Common::Thread::IsMainThread());
auto* slist = Common::SubsystemsList::Instance();
@@ -127,12 +127,21 @@ static void Init(const Config::ConfigOptions& cfg) {
slist->InitAll(true);
Config::Load(cfg);
slist->Add(log, {core, config});
slist->InitAll(true);
if (Common::File::IsFileExisting(param_json)) {
Loader::SystemContentLoadParamSfo(param_json);
if (const auto flexible_memory_size = Loader::SystemContentGetFlexibleMemorySize();
flexible_memory_size != 0) {
Libs::LibKernel::Memory::SetFlexibleMemorySize(flexible_memory_size);
}
}
slist->Add(audio, {core, log, pthread, memory});
slist->Add(controller, {core, log, config});
slist->Add(file_system, {core, log, pthread});
slist->Add(graphics, {core, log, pthread, memory, config, profiler, controller});
slist->Add(log, {core, config});
slist->Add(memory, {core, log});
slist->Add(network, {core, log, pthread});
slist->Add(profiler, {core, config});
@@ -180,7 +189,8 @@ void Run(const RunOptions& options) {
EXIT("ELF is required\n");
}
Init(options.config);
const auto param_json = options.app0_dir / "sce_sys" / "param.json";
Init(options.config, param_json);
ClearDebugTextureFolder();
@@ -192,15 +202,6 @@ void Run(const RunOptions& options) {
Libs::LibKernel::FileSystem::Mount(options.app0_dir, "/app0");
Libs::LibKernel::FileSystem::Mount(options.app0_dir, "/hostapp");
auto param_json = options.app0_dir / "sce_sys" / "param.json";
if (Common::File::IsFileExisting(param_json)) {
Loader::SystemContentLoadParamSfo(param_json);
if (auto flexible_memory_size = Loader::SystemContentGetFlexibleMemorySize();
flexible_memory_size != 0) {
Libs::LibKernel::Memory::SetFlexibleMemorySize(flexible_memory_size);
}
}
MountSandboxDirs();
auto* rt = Common::Singleton<Loader::RuntimeLinker>::Instance();
@@ -158,7 +158,7 @@ private:
void CheckBuffer() const { GetScheduler().CheckActive(); }
GpuResourceManager& GetGpuResources() const { return m_renderer.GetGpuResources(); }
RenderContext& m_renderer;
RenderContext& m_renderer;
HW::Context m_ctx;
HW::UserConfig m_ucfg;
HW::Shader m_sh_ctx;
@@ -170,9 +170,9 @@ private:
uint64_t m_dispatch_indirect_args_base_addr = 0;
uint32_t m_num_instances = 1;
uint32_t m_de_count = 0;
uint32_t m_ce_count = 0;
bool m_ce_complete = false;
uint32_t m_de_count = 0;
uint32_t m_ce_count = 0;
bool m_ce_complete = false;
bool m_readback_active = false;
uint32_t m_const_ram[0x3000] = {0};
@@ -1917,17 +1917,23 @@ KYTY_CP_OP_PARSER(CpOpCopyData) {
EXIT_NOT_IMPLEMENTED(cmd_id != KYTY_PM4(6, Pm4::IT_COPY_DATA, 0u));
const uint32_t control = buffer[0];
const uint32_t src_sel = ((control & 0xfu) << 1u) | ((control >> 30u) & 0x1u);
const uint32_t dst_sel = ((control >> 8u) & 0xfu) << 1u;
const uint8_t src_cache = static_cast<uint8_t>((control >> 13u) & 0x3u);
const uint8_t dst_cache = static_cast<uint8_t>((control >> 25u) & 0x3u);
const uint8_t write_confirm = static_cast<uint8_t>((control >> 20u) & 0x1u);
const uint32_t num_bytes = ((control >> 16u) & 0x1u) != 0 ? 8u : 4u;
const uint64_t src = buffer[1] | (static_cast<uint64_t>(buffer[2]) << 32u);
const uint64_t dst = buffer[3] | (static_cast<uint64_t>(buffer[4]) << 32u);
if (src_sel == (9u << 1u)) {
if (dst_sel != (2u << 1u) || dst == 0 || (dst & (num_bytes - 1u)) != 0) {
const uint32_t control = buffer[0];
const uint32_t src_sel = ((control & 0xfu) << 1u) | ((control >> 30u) & 0x1u);
const uint32_t dst_sel = ((control >> 8u) & 0xfu) << 1u;
const uint8_t src_cache = static_cast<uint8_t>((control >> 13u) & 0x3u);
const uint8_t dst_cache = static_cast<uint8_t>((control >> 25u) & 0x3u);
const uint8_t write_confirm = static_cast<uint8_t>((control >> 20u) & 0x1u);
const uint32_t num_bytes = ((control >> 16u) & 0x1u) != 0 ? 8u : 4u;
const uint64_t src = buffer[1] | (static_cast<uint64_t>(buffer[2]) << 32u);
const uint64_t dst = buffer[3] | (static_cast<uint64_t>(buffer[4]) << 32u);
uint32_t reference_clock_dst = 0;
switch (src_sel) {
case 9u: reference_clock_dst = 2u; break;
case 18u: reference_clock_dst = 4u; break;
default: break;
}
if (reference_clock_dst != 0) {
if (dst_sel != reference_clock_dst || dst == 0 || (dst & (num_bytes - 1u)) != 0) {
EXIT("unsupported reference-clock copyData, src_sel=0x%02" PRIx32
" dst_sel=0x%02" PRIx32 " dst=0x%016" PRIx64 " size=%u\n",
src_sel, dst_sel, dst, num_bytes);
@@ -3390,6 +3396,12 @@ void GraphicsInitJmpTablesCxIndirect() {
g_hw_ctx_indirect_func[Pm4::DB_COUNT_CONTROL] = [](KYTY_HW_CTX_INDIRECT_ARGS) {
HwCtxIgnoreDepthMetadataRegister(cmd_offset, value);
};
for (auto cmd_offset = Pm4::DB_SRESULTS_COMPARE_STATE0;
cmd_offset <= Pm4::DB_SRESULTS_COMPARE_STATE1; cmd_offset++) {
g_hw_ctx_indirect_func[cmd_offset] = [](KYTY_HW_CTX_INDIRECT_ARGS) {
HwCtxIgnoreDepthMetadataRegister(cmd_offset, value);
};
}
g_hw_ctx_indirect_func[Pm4::DB_RENDER_OVERRIDE] = [](KYTY_HW_CTX_INDIRECT_ARGS) {
HwCtxIgnoreDepthMetadataRegister(cmd_offset, value);
};
+3
View File
@@ -72,6 +72,7 @@ enum class ChannelLayout : uint32_t {
k32_32 = 11,
k16_16_16_16 = 12,
k32_32_32_32 = 14,
k5_6_5 = 16,
k5_5_5_1 = 17,
k4_4_4_4 = 19,
kBc1 = 35,
@@ -374,6 +375,8 @@ enum class BufferFormat : uint32_t {
k32_32_32_32UInt = 75,
k32_32_32_32SInt = 76,
k32_32_32_32Float = 77,
k8Srgb = 128,
k8_8Srgb = 129,
k8_8_8_8Srgb = 130,
k9_9_9_5Float = 132,
k5_6_5UNorm = 133,
+3
View File
@@ -39,6 +39,7 @@ constexpr FormatInfo kFormatInfo[] = {
{GpuEnumValue(BufferFormat::k16_16Float), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k11_11_10Float), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k10_10_10_2UNorm), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k10_10_10_2UInt), 4, 0, 4, true, true},
{GpuEnumValue(BufferFormat::k8_8_8_8UNorm), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k8_8_8_8SNorm), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k8_8_8_8UInt), 4, 0, 4, true, true},
@@ -57,6 +58,8 @@ constexpr FormatInfo kFormatInfo[] = {
{GpuEnumValue(BufferFormat::k32_32_32_32UInt), 16, 0, 16, true, true},
{GpuEnumValue(BufferFormat::k32_32_32_32SInt), 16, 0, 16, false, false},
{GpuEnumValue(BufferFormat::k32_32_32_32Float), 16, 0, 16, true, false},
{GpuEnumValue(BufferFormat::k8Srgb), 1, 0, 0, true, false},
{GpuEnumValue(BufferFormat::k8_8Srgb), 2, 0, 0, true, false},
{GpuEnumValue(BufferFormat::k8_8_8_8Srgb), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k9_9_9_5Float), 4, 0, 0, true, false},
{GpuEnumValue(BufferFormat::k5_6_5UNorm), 2, 0, 2, true, false},
+22 -76
View File
@@ -32,11 +32,10 @@
namespace Libs::Graphics {
static thread_local CommandProcessor* g_current_processor = nullptr;
static thread_local Pm4Execution* g_current_execution = nullptr;
static thread_local uint32_t g_submission_pause_depth = 0;
static thread_local bool g_gpu_mutex_owned = false;
static thread_local bool g_gpu_thread = false;
static thread_local CommandProcessor* g_current_processor = nullptr;
static thread_local Pm4Execution* g_current_execution = nullptr;
static thread_local bool g_gpu_mutex_owned = false;
static thread_local bool g_gpu_thread = false;
class GpuMutexLock final {
public:
@@ -98,8 +97,6 @@ public:
bool trigger_agc_interrupt_on_done);
void SubmitFlipPreparation(uint64_t request_id);
void Done();
void PauseSubmissions();
void ResumeSubmissions();
void Shutdown();
[[nodiscard]] bool IsStopping();
void SendCommand(Common::UniqueFunction<void>&& command);
@@ -410,8 +407,13 @@ void CommandProcessor::WriteData(uint32_t* dst, const uint32_t* src, uint32_t dw
const uint32_t increment = (write_control >> 16u) & 0x1u;
const uint32_t write_confirm = (write_control >> 20u) & 0x1u;
if (dst_sel != 0 && dst_sel != 2 && dst_sel != 4 && dst_sel != 5) {
EXIT("unsupported writeData destination selector 0x%02" PRIx32 "\n", dst_sel);
switch (dst_sel) {
case 0:
case 2:
case 4:
case 5:
case 6: break;
default: EXIT("unsupported writeData destination selector 0x%02" PRIx32 "\n", dst_sel);
}
EXIT_NOT_IMPLEMENTED(increment != 0);
@@ -691,26 +693,6 @@ bool GpuState::Process(Submission& submission) {
return complete;
}
void GpuState::PauseSubmissions() {
if (g_gpu_mutex_owned) {
EXIT("GPU submissions are already paused by this thread\n");
}
g_gpu_mutex_owned = true;
m_submission_mutex.Lock();
if (!IsGpuThread()) {
WaitLocked();
}
m_renderer.GetCommandScheduler().DrainPriorityOperations();
}
void GpuState::ResumeSubmissions() {
if (!g_gpu_mutex_owned) {
EXIT("GPU submissions resumed without an active pause\n");
}
m_submission_mutex.Unlock();
g_gpu_mutex_owned = false;
}
Pm4ProcessResult CommandProcessor::Process(Pm4Execution& execution, uint32_t* buffer,
uint32_t size_dw) {
KYTY_PROFILER_BLOCK("CommandProcessor::Process");
@@ -962,9 +944,8 @@ void CommandProcessor::DrawIndexOffset(uint32_t index_offset, uint32_t index_cou
auto* index_addr = reinterpret_cast<const void*>(
m_index_base_addr + static_cast<uint64_t>(index_offset) * index_size);
m_renderer.GetRenderExecutor().DrawIndex(m_submit_id, CurrentBuffer(),
m_index_type_and_size, index_count, index_addr,
flags, 1, m_num_instances);
m_renderer.GetRenderExecutor().DrawIndex(m_submit_id, CurrentBuffer(), m_index_type_and_size,
index_count, index_addr, flags, 1, m_num_instances);
}
void CommandProcessor::DrawIndirect(uint32_t data_offset, uint32_t draw_initiator, bool indexed) {
@@ -1190,8 +1171,8 @@ void CommandProcessor::DispatchDirect(uint32_t thread_group_x, uint32_t thread_g
}
}
m_renderer.GetRenderExecutor().DispatchDirect(
m_submit_id, CurrentBuffer(), thread_group_x, thread_group_y, thread_group_z, mode);
m_renderer.GetRenderExecutor().DispatchDirect(m_submit_id, CurrentBuffer(), thread_group_x,
thread_group_y, thread_group_z, mode);
}
constexpr uint32_t DispatchInitiatorUseThreadDimensions = 1u << 5u;
@@ -1237,16 +1218,16 @@ void CommandProcessor::DrawIndexAuto(uint32_t index_count, uint32_t flags,
uint32_t first_vertex, uint32_t first_instance) {
CheckBuffer();
m_renderer.GetRenderExecutor().DrawAuto(
m_submit_id, CurrentBuffer(), index_count, flags, render_target_slice_offset,
instance_count, first_vertex, first_instance);
m_renderer.GetRenderExecutor().DrawAuto(m_submit_id, CurrentBuffer(), index_count, flags,
render_target_slice_offset, instance_count,
first_vertex, first_instance);
}
void CommandProcessor::WaitFlipDone(uint32_t video_out_handle, uint32_t display_buffer_index) {
BufferFlush();
m_renderer.GetVideoOut().WaitFlipDone(static_cast<int>(video_out_handle),
static_cast<int>(display_buffer_index));
static_cast<int>(display_buffer_index));
}
template <typename T>
@@ -1317,8 +1298,8 @@ void CommandProcessor::WriteAtEndOfPipe(uint32_t cache_policy, uint32_t event_wr
if (eop_event_type == 0x2f && cache_action == 0x00 && event_index == 0x06) {
auto* dst = static_cast<uint32_t*>(dst_gpu_addr);
SynchronizeGpu();
Sync::ReadGds(m_renderer.GetBufferCache().GetGdsBuffer(), dst,
value & 0xffffu, value >> 16u);
Sync::ReadGds(m_renderer.GetBufferCache().GetGdsBuffer(), dst, value & 0xffffu,
value >> 16u);
Sync::WriteAtEndOfPipeGds32(m_submit_id, CurrentBuffer(), dst, value & 0xffffu,
value >> 16u);
return;
@@ -1486,8 +1467,7 @@ void CommandProcessor::EmitGlobalBarrier() {
barrier.srcStageMask = vk::PipelineStageFlagBits2::eAllCommands;
barrier.srcAccessMask = vk::AccessFlagBits2::eMemoryWrite;
barrier.dstStageMask = vk::PipelineStageFlagBits2::eAllCommands;
barrier.dstAccessMask =
vk::AccessFlagBits2::eMemoryRead | vk::AccessFlagBits2::eMemoryWrite;
barrier.dstAccessMask = vk::AccessFlagBits2::eMemoryRead | vk::AccessFlagBits2::eMemoryWrite;
vk::DependencyInfo dependency {};
dependency.memoryBarrierCount = 1;
@@ -1690,32 +1670,6 @@ int Gpu::GetFrameNum() const {
return m_state->GetFrameNum();
}
void Gpu::PauseSubmissions() {
m_state->PauseSubmissions();
}
void Gpu::ResumeSubmissions() {
m_state->ResumeSubmissions();
}
Gpu::SubmissionLock::SubmissionLock(Gpu& gpu): m_gpu(gpu) {
if (g_current_processor != nullptr || g_submission_pause_depth == UINT32_MAX) {
EXIT("cannot acquire GPU submission lock in the current state\n");
}
if (g_submission_pause_depth++ == 0) {
m_gpu.PauseSubmissions();
}
}
Gpu::SubmissionLock::~SubmissionLock() {
if (g_submission_pause_depth == 0) {
EXIT("GPU submission lock released without ownership\n");
}
if (--g_submission_pause_depth == 0) {
m_gpu.ResumeSubmissions();
}
}
bool Gpu::IsCommandProcessorThread() noexcept {
return g_current_processor != nullptr;
}
@@ -1724,12 +1678,4 @@ CommandProcessor* Gpu::CurrentCommandProcessor() noexcept {
return g_current_processor;
}
bool Gpu::SubmissionLockHeld() noexcept {
return g_submission_pause_depth != 0;
}
bool Gpu::MutexHeld() noexcept {
return g_gpu_mutex_owned;
}
} // namespace Libs::Graphics
-17
View File
@@ -35,25 +35,8 @@ public:
[[nodiscard]] static bool IsCommandProcessorThread() noexcept;
[[nodiscard]] static CommandProcessor* CurrentCommandProcessor() noexcept;
[[nodiscard]] static bool SubmissionLockHeld() noexcept;
[[nodiscard]] static bool MutexHeld() noexcept;
class SubmissionLock final {
public:
explicit SubmissionLock(Gpu& gpu);
~SubmissionLock();
KYTY_CLASS_NO_COPY(SubmissionLock);
private:
Gpu& m_gpu;
};
private:
friend class SubmissionLock;
void PauseSubmissions();
void ResumeSubmissions();
std::unique_ptr<GpuState> m_state;
};
} // namespace Libs::Graphics
+3 -1
View File
@@ -110,7 +110,6 @@ void DumpPm4PacketStream(Common::File* file, uint32_t* cmd_buffer, uint32_t star
auto* cmd = cmd_buffer + start_dw;
auto dw = num_dw;
while (dw != 0) {
EXIT_NOT_IMPLEMENTED(dw < 2);
EXIT_NOT_IMPLEMENTED(dw > num_dw);
auto cmd_id = *cmd++;
@@ -120,6 +119,9 @@ void DumpPm4PacketStream(Common::File* file, uint32_t* cmd_buffer, uint32_t star
uint32_t len = 0;
const auto packet_type = static_cast<PacketType>(cmd_id >> 30u);
// Type-2 packets are header-only padding; every other packet type requires a body.
EXIT_NOT_IMPLEMENTED(dw < 2 && packet_type != PacketType::Type2);
switch (packet_type) {
case PacketType::Type3: {
const bool sh_gx = (cmd_id & 0x2u) == 0;
+11 -11
View File
@@ -65,41 +65,41 @@ struct TileVolumeLayout {
};
bool TileGetBlockLayout(TileBlockFamily family, uint32_t bytes_per_element,
TileBlockLayout& layout);
TileBlockLayout& layout);
bool TileGetBlockOffset(const TileBlockLayout& layout, uint32_t x, uint32_t y, uint32_t z,
uint32_t& byte_offset);
uint32_t& byte_offset);
bool TileGetBlockXor(const TileBlockLayout& layout, uint32_t block_x, uint32_t block_y,
uint32_t& byte_offset);
uint32_t& byte_offset);
bool TileGetBlockXor(const TileBlockLayout& layout, uint32_t block_x, uint32_t block_y,
uint32_t block_z, uint32_t& byte_offset);
uint32_t block_z, uint32_t& byte_offset);
bool TileIsStandard256BTextureSupported(uint32_t format);
bool TileIsStandard4KBTextureSupported(uint32_t format);
bool TileIsStandard64KBTextureSupported(uint32_t format);
bool TileGetTextureVolumeLayout(uint32_t format, uint32_t width, uint32_t height, uint32_t depth,
uint32_t levels, uint32_t tile, TileVolumeLayout& layout);
uint32_t levels, uint32_t tile, TileVolumeLayout& layout);
bool TileGetHtileSize(uint32_t width, uint32_t height, TileSizeAlign& htile_size);
bool TileGetDepthSize(uint32_t width, uint32_t height, uint32_t pitch, uint32_t z_format,
uint32_t stencil_format, bool htile, TileSizeAlign& stencil_size,
TileSizeAlign& htile_size, TileSizeAlign& depth_size,
uint32_t stencil_format, bool htile, TileSizeAlign& stencil_size,
TileSizeAlign& htile_size, TileSizeAlign& depth_size,
uint32_t num_fragments_log2 = 0);
uint32_t TileGetRenderTargetPitch(uint32_t width, uint32_t bytes_per_element,
uint32_t num_fragments_log2 = 0);
uint32_t TileGetDepthPitch(uint32_t width, uint32_t bytes_per_element,
uint32_t num_fragments_log2 = 0);
bool TileGetRenderTargetSize(uint32_t width, uint32_t height, uint32_t pitch,
uint32_t bytes_per_element, TileSizeAlign& total_size,
uint32_t bytes_per_element, TileSizeAlign& total_size,
uint32_t num_fragments_log2 = 0);
bool TileGetRenderTargetMipLayout(uint32_t width, uint32_t height, uint32_t pitch,
uint32_t bytes_per_element, uint32_t levels,
TileSizeAlign& total_size, TileSizeOffset* level_sizes,
uint32_t bytes_per_element, uint32_t levels,
TileSizeAlign& total_size, TileSizeOffset* level_sizes,
TilePaddedSize* padded_size);
void TileGetTextureSize(uint32_t format, uint32_t width, uint32_t height, uint32_t pitch,
uint32_t levels, uint32_t tile, TileSizeAlign* total_size,
TileSizeOffset* level_sizes, TilePaddedSize* padded_size);
void TileGetTextureTotalSize(uint32_t format, uint32_t width, uint32_t height, uint32_t depth,
uint32_t pitch, uint32_t levels, uint32_t tile, bool volume_texture,
TileSizeAlign& total_size);
TileSizeAlign& total_size);
uint32_t TileGetTexturePitch(uint32_t format, uint32_t width, uint32_t levels, uint32_t tile);
} // namespace Libs::Graphics
+12 -12
View File
@@ -60,19 +60,19 @@ struct VulkanImage {
VulkanImage() = default;
KYTY_CLASS_NO_COPY(VulkanImage);
vk::Format format = vk::Format::eUndefined;
vk::ImageType image_type = vk::ImageType::e2D;
vk::Extent3D extent = {1, 1, 1};
uint32_t guest_pitch = 0;
uint32_t layers = 1;
uint32_t mip_levels = 1;
uint32_t samples = 1;
vk::ImageUsageFlags usage = {};
vk::ImageCreateFlags flags = {};
vk::Image image = nullptr;
VulkanImageState state;
vk::Format format = vk::Format::eUndefined;
vk::ImageType image_type = vk::ImageType::e2D;
vk::Extent3D extent = {1, 1, 1};
uint32_t guest_pitch = 0;
uint32_t layers = 1;
uint32_t mip_levels = 1;
uint32_t samples = 1;
vk::ImageUsageFlags usage = {};
vk::ImageCreateFlags flags = {};
vk::Image image = nullptr;
VulkanImageState state;
std::vector<VulkanImageState> subresource_states;
Graphics::VulkanMemory memory;
Graphics::VulkanMemory memory;
};
struct VulkanBuffer {
+1 -1
View File
@@ -30,7 +30,7 @@ bool IsAccessible(DWORD protect, HostMemoryAccess access) {
} // namespace
bool HostMemoryQueryRange(uint64_t addr, uint64_t requested_size, HostMemoryAccess access,
uint64_t& accessible_size) {
uint64_t& accessible_size) {
accessible_size = 0;
if (addr == 0 || requested_size == 0) {
return false;
+1 -1
View File
@@ -8,7 +8,7 @@ namespace Libs::Graphics {
enum class HostMemoryAccess { Read, Mapped };
bool HostMemoryQueryRange(uint64_t addr, uint64_t requested_size, HostMemoryAccess access,
uint64_t& accessible_size);
uint64_t& accessible_size);
bool HostMemoryQueryReadable(uint64_t addr, uint64_t requested_size, uint64_t& readable_size);
bool HostMemoryIsReadable(uint64_t addr);
bool HostMemoryRangeIsReadable(uint64_t addr, uint64_t size);
+9 -147
View File
@@ -4,25 +4,9 @@
namespace Libs::Graphics {
#if defined(KYTY_MEMORY_TRACKER_TESTS)
namespace {
std::atomic<MemoryTracker::UnmapContentionHook> g_unmap_contention_hook {nullptr};
}
void MemoryTracker::SetUnmapContentionHook(UnmapContentionHook hook) noexcept {
g_unmap_contention_hook.store(hook, std::memory_order_release);
}
#endif
static_assert(std::atomic<void*>::is_always_lock_free);
MemoryTracker::MemoryTracker(PageManager& page_manager, PageWatchMode gpu_watch_mode)
: m_page_manager(page_manager), m_gpu_watch_mode(gpu_watch_mode) {
switch (m_gpu_watch_mode) {
case PageWatchMode::Write:
case PageWatchMode::ReadWrite: break;
default: EXIT("unsupported memory tracker GPU page-watch mode\n");
}
MemoryTracker::MemoryTracker(PageManager& page_manager): m_page_manager(page_manager) {
m_regions = std::make_unique<std::atomic<RegionManager*>[]>(REGION_COUNT);
for (size_t i = 0; i < REGION_COUNT; i++) {
m_regions[i].store(nullptr, std::memory_order_relaxed);
@@ -31,6 +15,7 @@ MemoryTracker::MemoryTracker(PageManager& page_manager, PageWatchMode gpu_watch_
MemoryTracker::~MemoryTracker() = default;
#if KYTY_BUILD == KYTY_BUILD_DEBUG
void MemoryTracker::ValidateGpuDirtyPages(const RangeSet& dirty, uint64_t vaddr, uint64_t size,
const char* operation) const noexcept {
if (vaddr == 0 || size == 0 || size > UINT64_MAX - vaddr ||
@@ -68,6 +53,7 @@ void MemoryTracker::ValidateGpuDirtyOwnership(const RangeSet& dirty, uint64_t va
}
}
}
#endif
void MemoryTracker::ValidateRange(uint64_t vaddr, uint64_t size) {
if (vaddr == 0 || size == 0 || vaddr >= TRACKER_ADDRESS_SIZE ||
@@ -94,7 +80,6 @@ RegionManager* MemoryTracker::GetOrCreateRegion(uint64_t index) {
bool MemoryTracker::IsRegionCpuModified(uint64_t vaddr, uint64_t size) {
CheckNotInUploadCallback();
std::lock_guard access(m_access_mutex);
RequireMapped(vaddr, size);
return Iterate<true>(vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
return manager->IsModified<DirtySource::Cpu>(offset, bytes);
@@ -104,7 +89,6 @@ bool MemoryTracker::IsRegionCpuModified(uint64_t vaddr, uint64_t size) {
bool MemoryTracker::IsRegionGpuModified(uint64_t vaddr, uint64_t size) {
CheckNotInUploadCallback();
std::lock_guard access(m_access_mutex);
RequireMapped(vaddr, size);
return Iterate<false>(vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
return manager->IsModified<DirtySource::Gpu>(offset, bytes);
@@ -114,45 +98,31 @@ bool MemoryTracker::IsRegionGpuModified(uint64_t vaddr, uint64_t size) {
void MemoryTracker::MarkRegionAsCpuModified(uint64_t vaddr, uint64_t size) {
CheckNotInUploadCallback();
std::lock_guard access(m_access_mutex);
RequireMapped(vaddr, size);
Iterate<true>(vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
const auto changed =
manager->ChangeState<DirtySource::Cpu, true>(manager->GetCpuAddr() + offset, bytes);
manager->ApplyProtection(changed, false);
manager->ChangeState<DirtySource::Cpu, true>(manager->GetCpuAddr() + offset, bytes);
});
}
void MemoryTracker::MarkRegionAsGpuModified(uint64_t vaddr, uint64_t size) {
CheckNotInUploadCallback();
std::lock_guard access(m_access_mutex);
RequireMapped(vaddr, size);
Iterate<true>(vaddr, size, [this](RegionManager* manager, uint64_t offset, uint64_t bytes) {
Iterate<true>(vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
const auto changed =
manager->ChangeState<DirtySource::Gpu, true>(manager->GetCpuAddr() + offset, bytes);
manager->ApplyGpuProtection(changed, true, m_gpu_watch_mode);
manager->ChangeState<DirtySource::Gpu, true>(manager->GetCpuAddr() + offset, bytes);
});
}
void MemoryTracker::UnmarkRegionAsGpuModified(uint64_t vaddr, uint64_t size) {
CheckNotInUploadCallback();
std::lock_guard access(m_access_mutex);
RequireMapped(vaddr, size);
Iterate<true>(vaddr, size, [this](RegionManager* manager, uint64_t offset, uint64_t bytes) {
Iterate<false>(vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
if (!manager->IsFullyModified<DirtySource::Gpu>(offset, bytes)) {
EXIT("cannot clear partially GPU-dirty tracking range\n");
}
const auto changed =
manager->ChangeState<DirtySource::Gpu, false>(manager->GetCpuAddr() + offset, bytes);
manager->ApplyGpuProtection(changed, false, m_gpu_watch_mode);
manager->ChangeState<DirtySource::Gpu, false>(manager->GetCpuAddr() + offset, bytes);
});
}
void MemoryTracker::UntrackMemoryLocked(uint64_t vaddr, uint64_t size) {
RequireMapped(vaddr, size);
std::vector<RegionManager*> managers;
managers.reserve((vaddr % TRACKER_REGION_SIZE + size + TRACKER_REGION_SIZE - 1) /
TRACKER_REGION_SIZE);
@@ -171,10 +141,7 @@ void MemoryTracker::UntrackMemoryLocked(uint64_t vaddr, uint64_t size) {
EXIT("cannot untrack GPU-dirty memory\n");
}
Iterate<false>(vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
const auto changed =
manager->ChangeState<DirtySource::Cpu, true>(manager->GetCpuAddr() + offset, bytes);
manager->ApplyProtection(changed, false);
manager->Untrack(manager->GetCpuAddr() + offset, bytes);
manager->ChangeState<DirtySource::Cpu, true>(manager->GetCpuAddr() + offset, bytes);
});
locks.clear();
}
@@ -185,109 +152,4 @@ void MemoryTracker::UntrackMemory(uint64_t vaddr, uint64_t size) {
UntrackMemoryLocked(vaddr, size);
}
void MemoryTracker::UnmapMemory(uint64_t vaddr, uint64_t size) {
CheckNotInUploadCallback();
std::unique_lock access(m_access_mutex, std::try_to_lock);
if (!access.owns_lock()) {
#if defined(KYTY_MEMORY_TRACKER_TESTS)
if (const auto hook = g_unmap_contention_hook.load(std::memory_order_acquire);
hook != nullptr) {
hook();
}
#endif
access.lock();
}
UntrackMemoryLocked(vaddr, size);
m_page_manager.OnGpuUnmap(vaddr, size);
}
bool MemoryTracker::InvalidateRegion(uint64_t vaddr, uint64_t size, PageFaultPhase phase) noexcept {
switch (phase) {
case PageFaultPhase::Release: return true;
case PageFaultPhase::Invalidate: {
const auto action = BeginCpuFault(vaddr, size);
switch (action) {
case CpuFaultAction::Untracked: return false;
case CpuFaultAction::Continue: return true;
case CpuFaultAction::Download:
EXIT("generic region invalidation cannot download GPU-dirty memory\n");
}
}
case PageFaultPhase::Complete:
return CompleteCpuFault(vaddr, size, PageFaultAccess::Write, false);
}
EXIT("unsupported region invalidation phase\n");
}
bool MemoryTracker::InvalidateVirtualGpuWrite(PageFaultAccess access, uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept {
switch (phase) {
case PageFaultPhase::Release: return true;
case PageFaultPhase::Invalidate: {
const bool gpu_modified = Iterate<false>(
vaddr, size, [](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
return manager->IsModified<DirtySource::Gpu>(offset, bytes);
});
if (!gpu_modified) {
return false;
}
const auto action = BeginCpuFault(vaddr, size);
if (access != PageFaultAccess::Write || action != CpuFaultAction::Download) {
EXIT("virtual GPU write fault requires write access to GPU-dirty memory\n");
}
return true;
}
case PageFaultPhase::Complete: {
if (access != PageFaultAccess::Write) {
EXIT("virtual GPU write completion requires write access\n");
}
bool completed = false;
Iterate<false>(
vaddr, size, [&completed](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
if (completed) {
EXIT("virtual GPU write fault spans multiple tracked regions\n");
}
completed =
manager->CompleteVirtualGpuWrite(manager->GetCpuAddr() + offset, bytes);
});
return completed;
}
}
EXIT("unsupported virtual GPU write invalidation phase\n");
}
CpuFaultAction MemoryTracker::BeginCpuFault(uint64_t vaddr, uint64_t size,
PageFaultAccess access) noexcept {
CheckNotInUploadCallback();
CpuFaultAction action = CpuFaultAction::Untracked;
Iterate<false>(
vaddr, size, [&action, access](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
if (action != CpuFaultAction::Untracked) {
EXIT("CPU fault spans multiple tracked regions\n");
}
action = manager->BeginCpuFault(manager->GetCpuAddr() + offset, bytes, access);
});
return action;
}
bool MemoryTracker::CompleteCpuFault(uint64_t vaddr, uint64_t size, PageFaultAccess access,
bool downloaded) noexcept {
CheckNotInUploadCallback();
bool found = false;
Iterate<false>(
vaddr, size,
[&found, access, downloaded](RegionManager* manager, uint64_t offset, uint64_t bytes) {
std::scoped_lock lock(manager->lock);
if (found) {
EXIT("CPU fault completion spans multiple tracked regions\n");
}
found = manager->CompleteCpuFault(manager->GetCpuAddr() + offset, bytes, access,
downloaded);
});
return found;
}
} // namespace Libs::Graphics
+20 -50
View File
@@ -18,8 +18,7 @@ namespace Libs::Graphics {
class MemoryTracker final {
public:
explicit MemoryTracker(PageManager& page_manager,
PageWatchMode gpu_watch_mode = PageWatchMode::ReadWrite);
explicit MemoryTracker(PageManager& page_manager);
~MemoryTracker();
KYTY_CLASS_NO_COPY(MemoryTracker);
@@ -30,14 +29,6 @@ public:
void MarkRegionAsGpuModified(uint64_t vaddr, uint64_t size);
void UnmarkRegionAsGpuModified(uint64_t vaddr, uint64_t size);
void UntrackMemory(uint64_t vaddr, uint64_t size);
void UnmapMemory(uint64_t vaddr, uint64_t size);
[[nodiscard]] CpuFaultAction
BeginCpuFault(uint64_t vaddr, uint64_t size,
PageFaultAccess access = PageFaultAccess::Write) noexcept;
[[nodiscard]] bool CompleteCpuFault(uint64_t vaddr, uint64_t size, PageFaultAccess access,
bool downloaded) noexcept;
[[nodiscard]] bool InvalidateRegion(uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept;
template <typename Flush>
void InvalidateRegion(uint64_t vaddr, uint64_t size, Flush&& on_flush) {
static_assert(std::is_invocable_v<Flush&>);
@@ -64,9 +55,8 @@ public:
}
Iterate<false>(vaddr, size,
[](RegionManager* manager, uint64_t offset, uint64_t bytes) {
const auto changed = manager->ChangeState<DirtySource::Cpu, true>(
manager->ChangeState<DirtySource::Cpu, true>(
manager->GetCpuAddr() + offset, bytes);
manager->ApplyProtection(changed, false);
});
return false;
};
@@ -79,20 +69,22 @@ public:
EXIT("memory invalidation retained GPU-owned pages\n");
}
}
[[nodiscard]] bool InvalidateVirtualGpuWrite(PageFaultAccess access, uint64_t vaddr,
uint64_t size, PageFaultPhase phase) noexcept;
void ValidateGpuDirtyPages(const RangeSet& dirty, uint64_t vaddr, uint64_t size,
const char* operation) const noexcept;
#if KYTY_BUILD == KYTY_BUILD_DEBUG
void ValidateGpuDirtyPages(const RangeSet& dirty, uint64_t vaddr, uint64_t size,
const char* operation) const noexcept;
void ValidateGpuDirtyOwnership(const RangeSet& dirty, uint64_t vaddr, uint64_t size,
const char* operation);
#else
void ValidateGpuDirtyPages(const RangeSet&, uint64_t, uint64_t, const char*) const noexcept {}
void ValidateGpuDirtyOwnership(const RangeSet&, uint64_t, uint64_t, const char*) {}
#endif
template <bool clear, typename Preflight, typename Func>
void ForEachDownloadRange(uint64_t vaddr, uint64_t size, Preflight&& preflight, Func&& func) {
static_assert(std::is_nothrow_invocable_v<Preflight&, uint64_t, uint64_t>);
static_assert(std::is_nothrow_invocable_v<Func&, uint64_t, uint64_t>);
CheckNotInUploadCallback();
std::lock_guard access(m_access_mutex);
RequireMapped(vaddr, size);
std::lock_guard access(m_access_mutex);
std::vector<RegionManager*> managers;
Iterate<false>(vaddr, size, [&](RegionManager* manager, uint64_t, uint64_t) {
managers.push_back(manager);
@@ -104,9 +96,6 @@ public:
}
Iterate<false>(vaddr, size, [&](RegionManager* manager, uint64_t offset, uint64_t bytes) {
const auto address = manager->GetCpuAddr() + offset;
if (manager->HasPendingFault(address, bytes)) {
EXIT("GPU download synchronization raced a pending CPU fault\n");
}
manager->template ForEachModifiedRange<DirtySource::Gpu, false>(address, bytes,
preflight);
});
@@ -118,10 +107,8 @@ public:
Iterate<false>(vaddr, size,
[&](RegionManager* manager, uint64_t offset, uint64_t bytes) {
const auto address = manager->GetCpuAddr() + offset;
const auto changed =
manager->template ForEachModifiedRange<DirtySource::Gpu, true>(
address, bytes, [](uint64_t, uint64_t) noexcept {});
manager->ApplyGpuProtection(changed, false, m_gpu_watch_mode);
manager->template ForEachModifiedRange<DirtySource::Gpu, true>(
address, bytes, [](uint64_t, uint64_t) noexcept {});
});
}
}
@@ -132,11 +119,6 @@ public:
vaddr, size, [](uint64_t, uint64_t) noexcept {}, std::forward<Func>(func));
}
#if defined(KYTY_MEMORY_TRACKER_TESTS)
using UnmapContentionHook = void (*)() noexcept;
static void SetUnmapContentionHook(UnmapContentionHook hook) noexcept;
#endif
template <typename RangeFunc, typename UploadFunc>
void ForEachUploadRange(uint64_t vaddr, uint64_t size, bool is_written, RangeFunc&& range_func,
UploadFunc&& upload_func) {
@@ -144,12 +126,10 @@ public:
static_assert(std::is_nothrow_invocable_v<UploadFunc&>);
CheckNotInUploadCallback();
std::unique_lock access(m_access_mutex);
RequireMapped(vaddr, size);
Iterate<true>(vaddr, size, [](RegionManager*, uint64_t, uint64_t) {});
const auto* previous_upload_owner = std::exchange(s_upload_owner, this);
Iterate<false>(vaddr, size, [&](RegionManager* manager, uint64_t offset, uint64_t bytes) {
manager->lock.lock();
manager->Track(manager->GetCpuAddr() + offset, bytes);
manager->ForEachModifiedRange<DirtySource::Cpu, true>(manager->GetCpuAddr() + offset,
bytes, range_func);
if (!is_written) {
@@ -158,13 +138,12 @@ public:
});
upload_func();
if (is_written) {
Iterate<false>(
vaddr, size, [this](RegionManager* manager, uint64_t offset, uint64_t bytes) {
const auto changed = manager->template ChangeState<DirtySource::Gpu, true>(
manager->GetCpuAddr() + offset, bytes);
manager->ApplyGpuProtection(changed, true, m_gpu_watch_mode);
manager->lock.unlock();
});
Iterate<false>(vaddr, size,
[](RegionManager* manager, uint64_t offset, uint64_t bytes) {
manager->template ChangeState<DirtySource::Gpu, true>(
manager->GetCpuAddr() + offset, bytes);
manager->lock.unlock();
});
}
s_upload_owner = previous_upload_owner;
}
@@ -209,16 +188,8 @@ private:
return false;
}
static void ValidateRange(uint64_t vaddr, uint64_t size);
void UntrackMemoryLocked(uint64_t vaddr, uint64_t size);
void RequireMapped(uint64_t vaddr, uint64_t size) const {
ValidateRange(vaddr, size);
if (!m_page_manager.IsMapped(vaddr, size)) {
EXIT("memory tracker range [0x%llx, 0x%llx) is not mapped\n",
static_cast<unsigned long long>(vaddr),
static_cast<unsigned long long>(vaddr + size));
}
}
static void ValidateRange(uint64_t vaddr, uint64_t size);
void UntrackMemoryLocked(uint64_t vaddr, uint64_t size);
RegionManager* GetOrCreateRegion(uint64_t index);
std::unique_ptr<std::atomic<RegionManager*>[]> m_regions;
@@ -226,7 +197,6 @@ private:
std::mutex m_region_mutex;
std::mutex m_access_mutex;
PageManager& m_page_manager;
PageWatchMode m_gpu_watch_mode = PageWatchMode::ReadWrite;
};
} // namespace Libs::Graphics
File diff suppressed because it is too large Load Diff
+8 -36
View File
@@ -2,60 +2,32 @@
#define EMULATOR_SRC_GRAPHICS_HOST_GPU_PAGEMANAGER_H_
#include "common/common.h"
#include "graphics/host_gpu/rangeSet.h"
#include "graphics/host_gpu/regionDefinitions.h"
#include <memory>
#include <span>
#include <vector>
namespace Libs::Graphics {
enum class PageFaultAccess { Read, Write, Execute, Unknown };
enum class PageFaultPhase { Invalidate, Complete, Release };
enum class PageWatchMode { Write, ReadWrite };
enum class GpuAccess { Read, Write, ReadWrite };
using PageFaultHandler = bool (*)(void* context, PageFaultAccess access, uint64_t vaddr,
uint64_t size, PageFaultPhase phase) noexcept;
class PageManager final {
public:
class BackingWrite final {
public:
BackingWrite(PageManager& manager, uint64_t vaddr, uint64_t size) noexcept;
~BackingWrite();
KYTY_CLASS_NO_COPY(BackingWrite);
private:
PageManager& m_manager;
uint64_t m_vaddr = 0;
uint64_t m_size = 0;
};
PageManager(PageFaultHandler fault_handler, void* fault_context);
PageManager();
// The owner must stop all PageManager callers before destruction.
~PageManager();
KYTY_CLASS_NO_COPY(PageManager);
[[nodiscard]] uint64_t GetPageSize() const;
[[nodiscard]] bool IsTracked(uint64_t vaddr) const noexcept;
[[nodiscard]] bool IsMapped(uint64_t vaddr, uint64_t size) const noexcept;
[[nodiscard]] bool HasGpuAccess(uint64_t vaddr, uint64_t size, GpuAccess access) const noexcept;
void UpdatePageWatchers(bool track, uint64_t vaddr, uint64_t size,
PageWatchMode mode = PageWatchMode::Write);
void OnGpuMap(uint64_t vaddr, uint64_t size, GpuAccess access = GpuAccess::ReadWrite);
void OnGpuUnmap(uint64_t vaddr, uint64_t size, GpuAccess access = GpuAccess::ReadWrite);
[[nodiscard]] bool HandleFault(PageFaultAccess access, uint64_t fault_vaddr) noexcept;
[[nodiscard]] std::vector<std::unique_ptr<BackingWrite>>
ReserveBackingWrites(std::span<const RangeSet::Range> ranges);
template <bool track>
void UpdatePageWatchers(uint64_t vaddr, uint64_t size);
template <bool track, bool is_read = false>
void UpdatePageWatchersForRegion(uint64_t base_addr, RegionBits& mask);
void OnGpuMap(uint64_t vaddr, uint64_t size);
void OnGpuUnmap(uint64_t vaddr, uint64_t size);
private:
void BeginBackingWrite(uint64_t vaddr, uint64_t size) noexcept;
void EndBackingWrite(uint64_t vaddr, uint64_t size) noexcept;
struct Impl;
std::unique_ptr<Impl> m_impl;
};
+3 -3
View File
@@ -1,10 +1,9 @@
#ifndef EMULATOR_SRC_GRAPHICS_HOST_GPU_REGIONDEFINITIONS_H_
#define EMULATOR_SRC_GRAPHICS_HOST_GPU_REGIONDEFINITIONS_H_
#include "common/bitArray.h"
#include "common/common.h"
#include <bitset>
namespace Libs::Graphics {
constexpr uint64_t TRACKER_PAGE_SIZE = 4ull * 1024ull;
@@ -13,7 +12,8 @@ constexpr uint64_t TRACKER_ADDRESS_SIZE = 1ull << 40u;
constexpr size_t TRACKER_REGION_PAGES = TRACKER_REGION_SIZE / TRACKER_PAGE_SIZE;
enum class DirtySource { Cpu, Gpu };
using RegionBits = std::bitset<TRACKER_REGION_PAGES>;
using RegionBits = Common::BitArray<TRACKER_REGION_PAGES>;
static_assert(sizeof(RegionBits) == TRACKER_REGION_PAGES / 8);
} // namespace Libs::Graphics
+53 -202
View File
@@ -25,8 +25,6 @@
namespace Libs::Graphics {
enum class CpuFaultAction { Untracked, Continue, Download };
class TrackingSpinLock final {
public:
void lock() noexcept {
@@ -78,229 +76,93 @@ public:
if (m_cpu_addr % TRACKER_REGION_SIZE != 0) {
EXIT("invalid region tracking manager construction\n");
}
m_cpu_dirty.set();
m_writable.set();
m_cpu_dirty.Fill();
m_writable.Fill();
m_readable.Fill();
}
KYTY_CLASS_NO_COPY(RegionManager);
[[nodiscard]] uint64_t GetCpuAddr() const { return m_cpu_addr; }
void Track(uint64_t vaddr, uint64_t size) {
const auto [start, end] = GetPageRange(vaddr, size);
for (auto page = start; page < end; page++) {
m_tracked.set(page);
}
}
void Untrack(uint64_t vaddr, uint64_t size) {
const auto [start, end] = GetPageRange(vaddr, size);
for (auto page = start; page < end; page++) {
m_tracked.reset(page);
}
}
template <DirtySource source>
[[nodiscard]] bool IsModified(uint64_t offset, uint64_t size) const {
const auto [start, end] = GetPageRange(m_cpu_addr + offset, size);
const auto& bits = GetBits<source>();
for (auto page = start; page < end; page++) {
if (bits.test(page)) {
return true;
}
}
return false;
}
template <DirtySource source>
[[nodiscard]] bool IsFullyModified(uint64_t offset, uint64_t size) const {
const auto [start, end] = GetPageRange(m_cpu_addr + offset, size);
const auto& bits = GetBits<source>();
for (auto page = start; page < end; page++) {
if (!bits.test(page)) {
return false;
}
}
return true;
return RegionBits(bits, start, end).Any();
}
template <DirtySource source, bool enable>
RegionBits ChangeState(uint64_t vaddr, uint64_t size) {
void ChangeState(uint64_t vaddr, uint64_t size) {
const auto [start, end] = GetPageRange(vaddr, size);
if constexpr (source == DirtySource::Cpu && enable) {
for (auto page = start; page < end; page++) {
if (m_gpu_dirty.test(page) || m_fault_pending.test(page)) {
EXIT("CPU dirty state conflicts with GPU dirty or pending fault state\n");
}
if (RegionBits(m_gpu_dirty, start, end).Any()) {
EXIT("CPU dirty state conflicts with GPU dirty state\n");
}
}
if constexpr (source == DirtySource::Gpu && enable) {
for (auto page = start; page < end; page++) {
if (m_cpu_dirty.test(page) || m_fault_pending.test(page)) {
EXIT("GPU dirty state conflicts with CPU dirty or pending fault state\n");
}
if (RegionBits(m_cpu_dirty, start, end).Any()) {
EXIT("GPU dirty state conflicts with CPU dirty state\n");
}
}
auto& bits = GetBits<source>();
auto changed = bits;
for (auto page = start; page < end; page++) {
bits.set(page, enable);
auto& bits = GetBits<source>();
if constexpr (enable) {
bits.SetRange(start, end);
} else {
bits.UnsetRange(start, end);
}
changed ^= bits;
if constexpr (source == DirtySource::Cpu) {
changed = m_cpu_dirty ^ m_writable;
m_writable = m_cpu_dirty;
UpdateCpuProtection<!enable>();
} else {
UpdateGpuProtection<enable>();
}
return changed;
}
[[nodiscard]] CpuFaultAction BeginCpuFault(uint64_t vaddr, uint64_t size,
PageFaultAccess access = PageFaultAccess::Write) {
if (access != PageFaultAccess::Read && access != PageFaultAccess::Write) {
EXIT("unsupported CPU fault access while beginning ownership transfer\n");
}
const auto [start, end] = GetPageRange(vaddr, size);
const bool tracked = m_tracked.test(start);
for (auto page = start; page < end; page++) {
if (m_tracked.test(page) != tracked) {
EXIT("CPU fault spans mixed tracked and untracked pages\n");
}
if (m_fault_pending.test(page)) {
return CpuFaultAction::Untracked;
}
if (m_cpu_dirty.test(page) != m_writable.test(page) ||
(m_gpu_dirty.test(page) && (m_cpu_dirty.test(page) || m_writable.test(page)))) {
EXIT("inconsistent CPU fault page state\n");
}
}
if (!tracked) {
return CpuFaultAction::Untracked;
}
bool gpu_dirty = m_gpu_dirty.test(start);
bool writable = m_writable.test(start);
for (auto page = start + 1; page < end; page++) {
if (m_gpu_dirty.test(page) != gpu_dirty || m_writable.test(page) != writable) {
EXIT("CPU fault spans pages with incompatible dirty or writable state\n");
}
}
for (auto page = start; page < end; page++) {
if (!gpu_dirty && access == PageFaultAccess::Write) {
m_cpu_dirty.set(page);
m_writable.set(page);
}
m_fault_pending.set(page);
}
return gpu_dirty ? CpuFaultAction::Download : CpuFaultAction::Continue;
}
[[nodiscard]] bool CompleteCpuFault(uint64_t vaddr, uint64_t size, PageFaultAccess access,
bool downloaded) {
const auto [start, end] = GetPageRange(vaddr, size);
for (auto page = start; page < end; page++) {
if (!m_fault_pending.test(page)) {
return false;
}
}
for (auto page = start; page < end; page++) {
const bool gpu_dirty = m_gpu_dirty.test(page);
if (gpu_dirty != downloaded) {
EXIT("CPU fault download result disagrees with GPU dirty state\n");
}
if (gpu_dirty) {
m_gpu_dirty.reset(page);
switch (access) {
case PageFaultAccess::Read: break;
case PageFaultAccess::Write:
m_cpu_dirty.set(page);
m_writable.set(page);
break;
default: EXIT("unsupported CPU fault access after GPU download\n");
}
}
m_fault_pending.reset(page);
}
return true;
}
[[nodiscard]] bool HasPendingFault(uint64_t vaddr, uint64_t size) const {
const auto [start, end] = GetPageRange(vaddr, size);
for (auto page = start; page < end; page++) {
if (m_fault_pending.test(page)) {
return true;
}
}
return false;
}
[[nodiscard]] bool CompleteVirtualGpuWrite(uint64_t vaddr, uint64_t size) {
const auto [start, end] = GetPageRange(vaddr, size);
for (auto page = start; page < end; page++) {
if (!m_fault_pending.test(page)) {
return false;
}
if (!m_gpu_dirty.test(page)) {
EXIT("virtual GPU write completion found a non-GPU-dirty page\n");
}
}
for (auto page = start; page < end; page++) {
m_gpu_dirty.reset(page);
m_cpu_dirty.set(page);
m_writable.set(page);
m_fault_pending.reset(page);
}
return true;
}
template <DirtySource source, bool clear, typename Func>
RegionBits ForEachModifiedRange(uint64_t vaddr, uint64_t size, Func&& func) {
void ForEachModifiedRange(uint64_t vaddr, uint64_t size, Func&& func) {
const auto [start, end] = GetPageRange(vaddr, size);
auto mask = GetBits<source>();
if constexpr (source == DirtySource::Cpu) {
mask &= ~m_fault_pending;
}
for (auto page = 0u; page < start; page++) {
mask.reset(page);
}
for (auto page = end; page < TRACKER_REGION_PAGES; page++) {
mask.reset(page);
}
RegionBits mask(GetBits<source>(), start, end);
if constexpr (clear) {
auto& bits = GetBits<source>();
for (auto page = start; page < end; page++) {
if (mask.test(page)) {
bits.reset(page);
}
}
GetBits<source>().UnsetRange(start, end);
}
if constexpr (source == DirtySource::Cpu && clear) {
auto changed = m_cpu_dirty ^ m_writable;
m_writable = m_cpu_dirty;
ApplyProtection(changed, true);
UpdateCpuProtection<true>();
ForEachRange(mask, std::forward<Func>(func));
return changed;
return;
}
if constexpr (source == DirtySource::Gpu && clear) {
UpdateGpuProtection<false>();
}
ForEachRange(mask, std::forward<Func>(func));
if constexpr (clear) {
return mask;
}
return {};
}
void ApplyProtection(const RegionBits& changed, bool track) {
ForEachRange(changed, [this, track](uint64_t vaddr, uint64_t size) {
m_page_manager.UpdatePageWatchers(track, vaddr, size);
});
}
void ApplyGpuProtection(const RegionBits& changed, bool track, PageWatchMode mode) {
if (mode != PageWatchMode::Write && mode != PageWatchMode::ReadWrite) {
EXIT("unsupported GPU page-watch mode\n");
}
ForEachRange(changed, [this, track, mode](uint64_t vaddr, uint64_t size) {
m_page_manager.UpdatePageWatchers(track, vaddr, size, mode);
});
}
TrackingSpinLock lock;
private:
template <bool track>
void UpdateCpuProtection() {
auto mask = m_cpu_dirty ^ m_writable;
m_writable = m_cpu_dirty;
if (mask.None()) {
return;
}
m_page_manager.UpdatePageWatchersForRegion<track>(m_cpu_addr, mask);
}
template <bool track>
void UpdateGpuProtection() {
auto readable = ~m_gpu_dirty;
auto mask = readable ^ m_readable;
m_readable = readable;
if (mask.None()) {
return;
}
if constexpr (track) {
m_page_manager.UpdatePageWatchersForRegion<true, true>(m_cpu_addr, mask);
} else {
m_page_manager.UpdatePageWatchersForRegion<false, true>(m_cpu_addr, mask);
}
}
template <DirtySource source>
RegionBits& GetBits() {
if constexpr (source == DirtySource::Cpu) {
@@ -331,18 +193,8 @@ private:
template <typename Func>
void ForEachRange(const RegionBits& bits, Func&& func) const {
size_t page = 0;
while (page < TRACKER_REGION_PAGES) {
while (page < TRACKER_REGION_PAGES && !bits.test(page)) {
page++;
}
const auto start = page;
while (page < TRACKER_REGION_PAGES && bits.test(page)) {
page++;
}
if (start != page) {
func(m_cpu_addr + start * TRACKER_PAGE_SIZE, (page - start) * TRACKER_PAGE_SIZE);
}
for (const auto [start, end]: bits) {
func(m_cpu_addr + start * TRACKER_PAGE_SIZE, (end - start) * TRACKER_PAGE_SIZE);
}
}
@@ -351,8 +203,7 @@ private:
RegionBits m_cpu_dirty;
RegionBits m_gpu_dirty;
RegionBits m_writable;
RegionBits m_fault_pending;
RegionBits m_tracked;
RegionBits m_readable;
};
} // namespace Libs::Graphics
+68 -448
View File
@@ -123,30 +123,6 @@ struct BufferCache::RetiredBuffer {
std::shared_ptr<Buffer> owner;
};
struct BufferCache::FaultReadback {
PageFaultAccess access = PageFaultAccess::Unknown;
uint64_t vaddr = 0;
uint64_t size = 0;
std::vector<DownloadRange> ranges;
bool installed = false;
[[nodiscard]] bool Active() const noexcept { return !ranges.empty(); }
void Reset() {
access = PageFaultAccess::Unknown;
vaddr = 0;
size = 0;
installed = false;
ranges.clear();
}
};
struct BufferCache::PendingBackingPublication {
uint64_t address = 0;
uint64_t size = 0;
uint64_t tick = 0;
};
std::pair<uint64_t, uint64_t> BufferCache::DownloadEnvelope(const DownloadCopy& copy) {
if (copy.owner == nullptr || copy.size == 0 || copy.source_offset > copy.owner->Size() ||
copy.size > copy.owner->Size() - copy.source_offset) {
@@ -218,34 +194,26 @@ void BufferCache::QueueGarbageDownload(std::span<const DownloadCopy> copies, Ret
if (copies.empty()) {
return;
}
auto downloads = RecordDownloads(copies);
const auto tick = m_scheduler.CurrentTick();
BeginBackingPublication(retire.address, retire.size, tick);
m_scheduler.DeferOperation([this, downloads = std::move(downloads), retire = std::move(retire),
tick]() mutable {
PublishDownloads(downloads);
{
FaultSafeCacheLock lock(this, m_mutex);
if (m_memory_tracker.IsRegionGpuModified(retire.address, retire.size)) {
m_memory_tracker.ForEachDownloadRange<true>(
retire.address, retire.size,
[&](uint64_t address, uint64_t size) noexcept {
m_memory_tracker.ValidateGpuDirtyPages(m_gpu_modified_ranges, address, size,
"asynchronous garbage retirement");
},
[](uint64_t, uint64_t) noexcept {});
}
for (const auto& range: downloads) {
m_gpu_modified_ranges.Subtract(range.address, range.size);
}
if (m_memory_tracker.IsRegionGpuModified(retire.address, retire.size) ||
!m_gpu_modified_ranges.Intersections(retire.address, retire.size).empty()) {
EXIT("BufferCache: asynchronous garbage collection retained GPU ownership\n");
}
m_memory_tracker.UntrackMemory(retire.address, retire.size);
}
CompleteBackingPublication(retire.address, retire.size, tick);
});
auto downloads = RecordDownloads(copies);
m_scheduler.DeferOperation(
[this, downloads = std::move(downloads), retire = std::move(retire)]() mutable {
PublishDownloads(downloads);
{
FaultSafeCacheLock lock(this, m_mutex);
for (const auto& range: downloads) {
m_gpu_modified_ranges.Subtract(range.address, range.size);
}
// ForEachDownloadRange reports full tracker pages, and every exact GPU-owned
// interval on those pages was downloaded and removed. Clearing the original
// query therefore cannot orphan a dirty sibling on an edge page.
m_memory_tracker.UnmarkRegionAsGpuModified(retire.address, retire.size);
if (m_memory_tracker.IsRegionGpuModified(retire.address, retire.size) ||
!m_gpu_modified_ranges.Intersections(retire.address, retire.size).empty()) {
EXIT("BufferCache: asynchronous garbage collection retained GPU ownership\n");
}
m_memory_tracker.UntrackMemory(retire.address, retire.size);
}
});
}
BufferCache::BufferCache(GraphicContext& graphics, CommandScheduler& scheduler,
@@ -253,13 +221,12 @@ BufferCache::BufferCache(GraphicContext& graphics, CommandScheduler& scheduler,
ResourceMutex& resource_mutex)
: m_graphics(graphics), m_scheduler(scheduler),
m_gds_buffer(graphics, scheduler, MemoryUsage::Stream, 0, AllFlags, GdsBufferSize),
m_fault_readback(std::make_unique<FaultReadback>()), m_memory_tracker(page_manager),
m_memory_tracker(page_manager),
m_staging_buffer(graphics, scheduler, MemoryUsage::Upload, 512 * MiB),
m_stream_buffer(graphics, scheduler, MemoryUsage::Stream, 64 * MiB),
m_download_buffer(graphics, scheduler, MemoryUsage::Download, 32 * MiB),
m_device_buffer(graphics, scheduler, MemoryUsage::DeviceLocal, 128 * MiB),
m_page_manager(page_manager), m_texture_cache(texture_cache),
m_resource_mutex(resource_mutex) {
m_texture_cache(texture_cache), m_resource_mutex(resource_mutex) {
std::memset(m_gds_buffer.Mapped().data(), 0, static_cast<size_t>(m_gds_buffer.Size()));
m_gds_buffer.Flush(0, m_gds_buffer.Size());
if (!m_graphics.CanReportMemoryUsage()) {
@@ -277,15 +244,9 @@ BufferCache::BufferCache(GraphicContext& graphics, CommandScheduler& scheduler,
}
BufferCache::~BufferCache() {
if (m_fault_readback->Active()) {
EXIT("BufferCache: destroyed with an active fault readback\n");
}
if (!m_gpu_modified_ranges.Empty()) {
EXIT("BufferCache: destroyed with pending GPU-modified ranges\n");
}
if (!m_pending_backing_publications.empty()) {
EXIT("BufferCache: destroyed with pending backing publications\n");
}
for (const auto& [vaddr, cached]: m_buffers) {
(void)vaddr;
if (m_memory_tracker.IsRegionGpuModified(cached->vaddr, cached->size)) {
@@ -295,68 +256,6 @@ BufferCache::~BufferCache() {
m_buffers.clear();
}
bool BufferCache::SynchronizeBacking(uint64_t vaddr, uint64_t size) {
bool waited = false;
for (;;) {
uint64_t tick = 0;
const auto page_begin = vaddr & ~(TRACKER_PAGE_SIZE - 1);
const auto page_end = (vaddr + size + TRACKER_PAGE_SIZE - 1) & ~(TRACKER_PAGE_SIZE - 1);
CacheRange affected {.address = page_begin, .size = page_end - page_begin};
{
FaultSafeCacheLock lock(this, m_mutex);
bool changed = true;
while (changed) {
changed = false;
for (const auto& [address, cached]: m_buffers) {
const CacheRange previous = affected;
if (ResolveOverlap(affected, {address, cached->size}) &&
(previous.address != affected.address || previous.size != affected.size)) {
changed = true;
}
}
}
}
{
std::lock_guard lock(m_publication_mutex);
for (const auto& publication: m_pending_backing_publications) {
if (publication.address < affected.address + affected.size &&
affected.address < publication.address + publication.size) {
tick = std::max(tick, publication.tick);
}
}
}
if (tick == 0) {
return waited;
}
waited = true;
m_scheduler.Wait(tick);
m_scheduler.WaitPriorityOperations(tick);
}
}
void BufferCache::RefreshInvalidatedRanges(CommandBuffer& command, CachedBuffer& cached,
uint64_t vaddr, uint64_t size, bool upload) {
const auto invalidated = m_image_invalidated_ranges.Intersections(vaddr, size);
if (upload) {
std::array<uint8_t, 64 * 1024> bytes;
for (const auto& range: invalidated) {
for (uint64_t copied = 0; copied < range.size;) {
const auto chunk = std::min<uint64_t>(range.size - copied, bytes.size());
if (!Libs::LibKernel::Memory::TryReadBacking(range.address + copied, bytes.data(),
chunk)) {
EXIT("BufferCache: failed to refresh an invalidated image alias\n");
}
Upload(command, *cached.buffer, cached.buffer->Offset(range.address + copied),
bytes.data(), chunk);
copied += chunk;
}
}
}
if (!invalidated.empty()) {
m_image_invalidated_ranges.Subtract(vaddr, size);
}
}
StreamBuffer& BufferCache::GetUtilityBuffer(MemoryUsage usage) noexcept {
switch (usage) {
case MemoryUsage::Upload: return m_staging_buffer;
@@ -385,7 +284,6 @@ void BufferCache::InvalidateMemory(uint64_t vaddr, uint64_t size) {
size > TRACKER_ADDRESS_SIZE - vaddr) {
EXIT("BufferCache: invalid memory-invalidation range\n");
}
(void)SynchronizeBacking(vaddr, size);
if (!HasPageOverlap(vaddr, size)) {
return;
}
@@ -394,7 +292,6 @@ void BufferCache::InvalidateMemory(uint64_t vaddr, uint64_t size) {
}
void BufferCache::ReadMemory(uint64_t vaddr, uint64_t size) {
(void)SynchronizeBacking(vaddr, size);
std::vector<DownloadCopy> copies;
{
FaultSafeCacheLock lock(this, m_mutex);
@@ -434,117 +331,21 @@ void BufferCache::ReadMemory(uint64_t vaddr, uint64_t size) {
PublishDownloads(downloads);
{
FaultSafeCacheLock lock(this, m_mutex);
m_memory_tracker.ForEachDownloadRange<true>(
vaddr, size,
[&](uint64_t address, uint64_t bytes) noexcept {
m_memory_tracker.ValidateGpuDirtyPages(m_gpu_modified_ranges, address, bytes,
"memory invalidation completion");
},
[](uint64_t, uint64_t) noexcept {});
for (const auto& range: downloads) {
m_gpu_modified_ranges.Subtract(range.address, range.size);
}
// The enumeration above covered whole dirty pages and every exact interval on them.
m_memory_tracker.UnmarkRegionAsGpuModified(vaddr, size);
}
}
bool BufferCache::InvalidateMemory(PageFaultAccess access, uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept {
const auto page = vaddr & ~(TRACKER_PAGE_SIZE - 1);
if (size == 0 || size > page + TRACKER_PAGE_SIZE - vaddr) {
EXIT("BufferCache: invalid page-fault range\n");
}
if (phase == PageFaultPhase::Complete) {
FaultSafeCacheLock lock(this, m_mutex);
auto& fault = *m_fault_readback;
if (!fault.Active()) {
return m_memory_tracker.CompleteCpuFault(vaddr, size, access, false);
}
if (fault.access != access || fault.vaddr != vaddr || fault.size != size ||
fault.installed) {
EXIT("BufferCache: mismatched fault readback completion\n");
}
PublishDownloads(fault.ranges);
if (!m_memory_tracker.CompleteCpuFault(vaddr, size, access, true)) {
EXIT("BufferCache: failed to complete downloaded CPU fault\n");
}
fault.installed = true;
return true;
}
if (phase == PageFaultPhase::Release) {
FaultSafeCacheLock lock(this, m_mutex);
auto& fault = *m_fault_readback;
if (fault.Active()) {
if (fault.access != access || fault.vaddr != vaddr || fault.size != size ||
!fault.installed) {
EXIT("BufferCache: mismatched fault readback release\n");
}
for (const auto& range: fault.ranges) {
m_gpu_modified_ranges.Subtract(range.address, range.size);
}
fault.Reset();
}
return true;
}
if (phase != PageFaultPhase::Invalidate) {
EXIT("BufferCache: unsupported page-fault phase\n");
}
const auto action = m_memory_tracker.BeginCpuFault(vaddr, size, access);
if (action != CpuFaultAction::Download) {
return action == CpuFaultAction::Continue;
}
auto& fault = *m_fault_readback;
std::vector<DownloadCopy> copies;
{
FaultSafeCacheLock lock(this, m_mutex);
if (fault.Active()) {
EXIT("BufferCache: nested fault readback\n");
}
fault.access = access;
fault.vaddr = vaddr;
fault.size = size;
m_gpu_modified_ranges.ForEachIntersection(
page, TRACKER_PAGE_SIZE, [&](RangeSet::Range range) {
auto owner = m_buffers.upper_bound(range.address);
if (owner == m_buffers.begin()) {
EXIT("BufferCache: fault readback has no buffer owner\n");
}
--owner;
auto& cached = *owner->second;
if (!cached.buffer->IsInBounds(range.address, range.size)) {
EXIT("BufferCache: fault readback is outside its buffer owner\n");
}
copies.push_back({cached.buffer, cached.buffer->Offset(range.address),
range.address, range.size});
});
if (copies.empty()) {
EXIT("BufferCache: GPU-dirty fault page has no dirty byte ranges\n");
}
}
fault.ranges = RecordDownloads(copies);
if (!fault.Active()) {
EXIT("BufferCache: GPU-dirty fault page has no dirty byte ranges\n");
}
m_scheduler.FinishCurrent();
return true;
}
void BufferCache::UnmapMemory(uint64_t vaddr, uint64_t size) {
if (vaddr == 0 || size == 0 || size > UINT64_MAX - vaddr) {
EXIT("BufferCache: invalid unmap range\n");
}
(void)SynchronizeBacking(vaddr, size);
std::vector<DownloadCopy> copies;
std::vector<RangeSet::Range> dirty_ranges;
std::vector<std::pair<uint64_t, uint64_t>> modified_buffers;
std::vector<std::unique_ptr<PageManager::BackingWrite>> backing_writes;
std::vector<std::pair<uint64_t, uint64_t>> retired_buffers;
std::vector<DownloadCopy> copies;
std::vector<std::pair<uint64_t, uint64_t>> modified_buffers;
std::vector<std::pair<uint64_t, uint64_t>> retired_buffers;
{
FaultSafeCacheLock lock(this, m_mutex);
for (const auto& [begin, cached]: m_buffers) {
@@ -561,12 +362,8 @@ void BufferCache::UnmapMemory(uint64_t vaddr, uint64_t size) {
if (dirty.empty()) {
EXIT("BufferCache: GPU-modified buffer has no dirty ranges\n");
}
dirty_ranges.insert(dirty_ranges.end(), dirty.begin(), dirty.end());
modified_buffers.emplace_back(begin, cached->size);
}
if (!dirty_ranges.empty()) {
backing_writes = m_page_manager.ReserveBackingWrites(dirty_ranges);
}
for (const auto& [begin, bytes]: modified_buffers) {
auto owner = m_buffers.find(begin);
if (owner == m_buffers.end() || owner->second->size != bytes) {
@@ -596,24 +393,11 @@ void BufferCache::UnmapMemory(uint64_t vaddr, uint64_t size) {
// command stream before removing such backing.
m_scheduler.FinishCurrent();
}
backing_writes.clear();
{
FaultSafeCacheLock lock(this, m_mutex);
for (const auto& [begin, bytes]: modified_buffers) {
if (!m_memory_tracker.IsRegionGpuModified(begin, bytes)) {
continue;
}
m_memory_tracker.ForEachDownloadRange<true>(
begin, bytes,
[&](uint64_t address, uint64_t download_size) noexcept {
m_memory_tracker.ValidateGpuDirtyPages(m_gpu_modified_ranges, address,
download_size, "unmap retirement");
},
[](uint64_t, uint64_t) noexcept {});
}
for (const auto& [begin, bytes]: modified_buffers) {
m_gpu_modified_ranges.Subtract(begin, bytes);
m_memory_tracker.UnmarkRegionAsGpuModified(begin, bytes);
}
for (const auto& [begin, bytes]: retired_buffers) {
m_memory_tracker.MarkRegionAsCpuModified(begin, bytes);
@@ -621,7 +405,6 @@ void BufferCache::UnmapMemory(uint64_t vaddr, uint64_t size) {
if (!m_gpu_modified_ranges.Intersections(vaddr, size).empty()) {
EXIT("BufferCache: unmap retained dirty byte ranges\n");
}
m_image_invalidated_ranges.Subtract(vaddr, size);
m_memory_tracker.UntrackMemory(vaddr, size);
for (auto it = m_buffers.begin(); it != m_buffers.end();) {
if (vaddr < it->first + it->second->size && it->first < vaddr + size) {
@@ -714,22 +497,30 @@ BufferBinding BufferCache::ObtainBuffer(CommandBuffer& command, uint64_t vaddr,
if (command.IsInvalid() || command.IsExecute()) {
EXIT("BufferCache: buffer request requires a recording command buffer\n");
}
ValidateGpuAccess(vaddr, size, is_read, is_written);
std::lock_guard transaction(m_resource_mutex);
(void)SynchronizeBacking(vaddr, size);
if (is_read && !is_written && size <= CACHING_PAGE_SIZE &&
!m_memory_tracker.IsRegionGpuModified(vaddr, size) &&
m_memory_tracker.IsRegionCpuModified(vaddr, size)) {
std::vector<uint8_t> data(size);
if (Libs::LibKernel::Memory::TryReadBacking(vaddr, data.data(), size)) {
return UploadTransient(data.data(), size, 16);
const auto alignment = std::max<uint64_t>(
m_graphics.physical_device_properties.limits.minUniformBufferOffsetAlignment, 1);
if (auto [mapped, offset] = m_stream_buffer.Map(size, alignment, false);
mapped != nullptr) {
if (Libs::LibKernel::Memory::TryReadBacking(vaddr, mapped, size)) {
m_stream_buffer.Commit();
return {{}, m_stream_buffer.Handle(), offset};
}
} else {
auto owner = std::make_shared<Buffer>(m_graphics, m_scheduler, MemoryUsage::Upload, 0,
AllFlags, size);
if (Libs::LibKernel::Memory::TryReadBacking(vaddr, owner->Mapped().data(), size)) {
owner->Flush(0, size);
return {owner, owner->Handle(), 0};
}
}
}
if (is_formatted && is_read && !is_written) {
(void)m_texture_cache.SynchronizeImageToBuffer(vaddr, size);
} else if (is_formatted && is_written) {
if (is_formatted && is_written) {
(void)m_texture_cache.InvalidateMemoryFromGPU(vaddr, size, true);
}
@@ -745,10 +536,12 @@ BufferBinding BufferCache::ObtainBuffer(CommandBuffer& command, uint64_t vaddr,
reinterpret_cast<const void*>(address), bytes);
}
});
RefreshInvalidatedRanges(command, cached, vaddr, size, is_read);
if (is_written) {
m_gpu_modified_ranges.Add(vaddr, size);
}
if (is_formatted && is_read && !is_written) {
(void)SynchronizeBufferFromImage(*cached.buffer, vaddr, size);
}
return {cached.buffer, cached.buffer->Handle(), cached.buffer->Offset(vaddr)};
}
@@ -773,7 +566,6 @@ ImageBufferSource BufferCache::ObtainBufferForImage(uint64_t vaddr, uint64_t siz
size > TRACKER_ADDRESS_SIZE - vaddr) {
EXIT("BufferCache: invalid image source\n");
}
(void)SynchronizeBacking(vaddr, size);
auto find_owner = [&]() {
auto owner = m_buffers.upper_bound(vaddr);
if (owner == m_buffers.begin()) {
@@ -788,13 +580,12 @@ ImageBufferSource BufferCache::ObtainBufferForImage(uint64_t vaddr, uint64_t siz
const bool cpu_modified = m_memory_tracker.IsRegionCpuModified(vaddr, size);
const bool gpu_modified = m_memory_tracker.IsRegionGpuModified(vaddr, size);
const auto dirty = m_gpu_modified_ranges.Intersections(vaddr, size);
const bool invalidated = !m_image_invalidated_ranges.Intersections(vaddr, size).empty();
const bool requested_gpu_owned = !dirty.empty();
const bool has_dirty_buffer_source = !dirty.empty();
m_memory_tracker.ValidateGpuDirtyOwnership(m_gpu_modified_ranges, vaddr, size,
"image source");
auto owner = find_owner();
if (requested_gpu_owned && owner == m_buffers.end()) {
if (has_dirty_buffer_source && owner == m_buffers.end()) {
CacheRange merged {.address = AlignDown(vaddr),
.size = AlignUp(vaddr + size) - AlignDown(vaddr)};
using Iterator = decltype(m_buffers.begin());
@@ -844,42 +635,32 @@ ImageBufferSource BufferCache::ObtainBufferForImage(uint64_t vaddr, uint64_t siz
EXIT("BufferCache: merged image source does not contain the requested range\n");
}
}
if (owner != m_buffers.end() && !cpu_modified && !invalidated &&
(!gpu_modified || requested_gpu_owned)) {
DiscardGpuDirtyBytesLocked(vaddr, size, "image source transfer");
if (owner != m_buffers.end() && !cpu_modified &&
(!gpu_modified || has_dirty_buffer_source)) {
owner->second->tick_accessed_last = m_gc_tick;
return {owner->second->buffer.get(), owner->second->buffer->Offset(vaddr),
requested_gpu_owned};
return {owner->second->buffer.get(), owner->second->buffer->Offset(vaddr)};
}
if (requested_gpu_owned && owner == m_buffers.end()) {
if (has_dirty_buffer_source && owner == m_buffers.end()) {
EXIT("BufferCache: GPU-dirty image source could not resolve its native owner\n");
}
}
// Direct-memory backing remains readable while PageManager protects the guest mapping. The
// fallback exists for plain host mappings used by standalone renderer tests and is deliberately
// performed outside the cache lock so a page fault cannot recurse into BufferCache.
const auto stage_address = vaddr & ~(TRACKER_PAGE_SIZE - 1);
const auto stage_end = (vaddr + size + TRACKER_PAGE_SIZE - 1) & ~(TRACKER_PAGE_SIZE - 1);
const auto stage_size = stage_end - stage_address;
(void)SynchronizeBacking(stage_address, stage_size);
std::vector<uint8_t> bytes(stage_size);
if (!Libs::LibKernel::Memory::TryReadBacking(stage_address, bytes.data(), stage_size)) {
auto [staging, stage_offset] = m_staging_buffer.Map(size, 16);
if (staging == nullptr || !Libs::LibKernel::Memory::TryReadBacking(vaddr, staging, size)) {
EXIT("BufferCache: failed to read mapped guest image backing\n");
}
m_staging_buffer.Commit();
FaultSafeCacheLock lock(this, m_mutex);
const auto dirty = m_gpu_modified_ranges.Intersections(vaddr, size);
const bool invalidated = !m_image_invalidated_ranges.Intersections(vaddr, size).empty();
const bool requested_gpu_owned = !dirty.empty();
auto owner = find_owner();
if (requested_gpu_owned && owner == m_buffers.end()) {
const auto dirty = m_gpu_modified_ranges.Intersections(vaddr, size);
const bool has_dirty_buffer_source = !dirty.empty();
auto owner = find_owner();
if (has_dirty_buffer_source && owner == m_buffers.end()) {
EXIT("BufferCache: GPU-dirty image source lost its native owner\n");
}
const auto stage_offset = m_staging_buffer.Copy(bytes.data(), stage_size, 16);
if (owner == m_buffers.end() || invalidated ||
(m_memory_tracker.IsRegionGpuModified(vaddr, size) && !requested_gpu_owned)) {
return {&m_staging_buffer, stage_offset + vaddr - stage_address, false};
if (owner == m_buffers.end() ||
(m_memory_tracker.IsRegionGpuModified(vaddr, size) && !has_dirty_buffer_source)) {
return {&m_staging_buffer, stage_offset};
}
auto& cached = *owner->second;
@@ -893,42 +674,17 @@ ImageBufferSource BufferCache::ObtainBufferForImage(uint64_t vaddr, uint64_t siz
[&]() noexcept {
for (const auto& [address, upload_size]: uploads) {
cached.buffer->CopyFrom(
m_scheduler.Current(), m_staging_buffer, stage_offset + address - stage_address,
m_scheduler.Current(), m_staging_buffer, stage_offset + address - vaddr,
cached.buffer->Offset(address), upload_size, vk::AccessFlagBits::eHostWrite);
}
});
DiscardGpuDirtyBytesLocked(vaddr, size, "staged image source transfer");
return {cached.buffer.get(), cached.buffer->Offset(vaddr), requested_gpu_owned};
}
void BufferCache::DiscardGpuDirtyBytesLocked(uint64_t vaddr, uint64_t size, const char* operation) {
m_memory_tracker.ValidateGpuDirtyOwnership(m_gpu_modified_ranges, vaddr, size, operation);
m_gpu_modified_ranges.Subtract(vaddr, size);
const auto page_begin = vaddr & ~(TRACKER_PAGE_SIZE - 1);
const auto page_end = (vaddr + size + TRACKER_PAGE_SIZE - 1) & ~(TRACKER_PAGE_SIZE - 1);
for (auto page = page_begin; page < page_end; page += TRACKER_PAGE_SIZE) {
if (m_gpu_modified_ranges.Intersections(page, TRACKER_PAGE_SIZE).empty() &&
m_memory_tracker.IsRegionGpuModified(page, TRACKER_PAGE_SIZE)) {
m_memory_tracker.UnmarkRegionAsGpuModified(page, TRACKER_PAGE_SIZE);
}
}
m_memory_tracker.ValidateGpuDirtyOwnership(m_gpu_modified_ranges, vaddr, size, operation);
}
void BufferCache::DiscardGpuDirtyBytes(uint64_t vaddr, uint64_t size) {
if (vaddr == 0 || size == 0 || vaddr >= TRACKER_ADDRESS_SIZE ||
size > TRACKER_ADDRESS_SIZE - vaddr) {
EXIT("BufferCache: invalid dirty-byte discard range\n");
}
FaultSafeCacheLock lock(this, m_mutex);
DiscardGpuDirtyBytesLocked(vaddr, size, "image output supersession");
return {cached.buffer.get(), cached.buffer->Offset(vaddr)};
}
void BufferCache::WriteHostMemory(uint64_t vaddr, std::span<const uint8_t> data) {
if (vaddr == 0 || data.empty() || data.size() > UINT64_MAX - vaddr) {
EXIT("BufferCache: invalid host DMA write\n");
}
(void)SynchronizeBacking(vaddr, data.size());
Libs::LibKernel::Memory::WriteBacking(vaddr, data.data(), data.size());
FaultSafeCacheLock lock(this, m_mutex);
@@ -944,45 +700,6 @@ void BufferCache::WriteHostMemory(uint64_t vaddr, std::span<const uint8_t> data)
data.data() + begin - vaddr, range_end - begin);
cached->tick_accessed_last = m_gc_tick;
}
m_image_invalidated_ranges.Subtract(vaddr, data.size());
}
std::pair<std::shared_ptr<Buffer>, uint64_t> BufferCache::ObtainBufferForImageWrite(uint64_t vaddr,
uint64_t size) {
if (vaddr == 0 || size == 0 || vaddr >= TRACKER_ADDRESS_SIZE ||
size > TRACKER_ADDRESS_SIZE - vaddr) {
EXIT("BufferCache: invalid image destination\n");
}
const auto stage_address = vaddr & ~(TRACKER_PAGE_SIZE - 1);
const auto stage_end = (vaddr + size + TRACKER_PAGE_SIZE - 1) & ~(TRACKER_PAGE_SIZE - 1);
const auto stage_size = stage_end - stage_address;
(void)SynchronizeBacking(stage_address, stage_size);
std::vector<uint8_t> bytes(stage_size);
if (!Libs::LibKernel::Memory::TryReadBacking(stage_address, bytes.data(), stage_size)) {
EXIT("BufferCache: failed to preserve guest bytes around an image mirror\n");
}
FaultSafeCacheLock lock(this, m_mutex);
auto& cached = GetOrCreateBuffer(m_scheduler.Current(), vaddr, size);
m_memory_tracker.ValidateGpuDirtyOwnership(m_gpu_modified_ranges, vaddr, size,
"image destination");
if (!m_gpu_modified_ranges.Intersections(vaddr, size).empty()) {
EXIT("BufferCache: image destination aliases GPU-owned buffer bytes\n");
}
const auto stage_offset = m_staging_buffer.Copy(bytes.data(), stage_size, 16);
std::vector<std::pair<uint64_t, uint64_t>> uploads;
m_memory_tracker.ForEachUploadRange(
vaddr, size, false,
[&](uint64_t address, uint64_t upload_size) noexcept {
uploads.emplace_back(address, upload_size);
},
[&]() noexcept {
for (const auto& [address, upload_size]: uploads) {
cached.buffer->CopyFrom(
m_scheduler.Current(), m_staging_buffer, stage_offset + address - stage_address,
cached.buffer->Offset(address), upload_size, vk::AccessFlagBits::eHostWrite);
}
});
return {cached.buffer, cached.buffer->Offset(vaddr)};
}
void BufferCache::FillBuffer(uint64_t vaddr, uint64_t size, uint32_t value, bool is_gds) {
@@ -999,7 +716,6 @@ void BufferCache::FillBuffer(uint64_t vaddr, uint64_t size, uint32_t value, bool
if (vaddr == 0) {
EXIT("BufferCache: invalid fill memory address\n");
}
ValidateGpuAccess(vaddr, size, false, true);
(void)m_texture_cache.ClearMeta(vaddr);
{
std::lock_guard transaction(m_resource_mutex);
@@ -1041,25 +757,12 @@ void BufferCache::CopyBuffer(uint64_t dst_vaddr, uint64_t src_vaddr, uint64_t si
(src_gds && (src_vaddr > m_gds_buffer.Size() || size > m_gds_buffer.Size() - src_vaddr))) {
EXIT("BufferCache: invalid or overlapping copy range\n");
}
if (src_memory) {
ValidateGpuAccess(src_vaddr, size, true, false);
}
if (dst_memory) {
ValidateGpuAccess(dst_vaddr, size, false, true);
}
if (src_memory || dst_memory) {
std::lock_guard transaction(m_resource_mutex);
if (src_memory) {
(void)SynchronizeBacking(src_vaddr, size);
}
const auto src_region =
const auto src_region =
src_memory ? m_texture_cache.QueryRegion(src_vaddr, size) : TextureCache::RegionInfo {};
const auto dst_region =
dst_memory ? m_texture_cache.QueryRegion(dst_vaddr, size) : TextureCache::RegionInfo {};
if (src_memory && src_region.gpu_image_bytes &&
!m_texture_cache.SynchronizeImageToBuffer(src_vaddr, size)) {
EXIT("BufferCache: GPU copy source image could not be synchronized\n");
}
if (src_memory && dst_memory && !HasGpuDirtyBytes(src_vaddr, size) &&
!HasGpuDirtyBytes(dst_vaddr, size) && !src_region.gpu_image_bytes &&
!dst_region.gpu_image_bytes) {
@@ -1081,7 +784,7 @@ void BufferCache::CopyBuffer(uint64_t dst_vaddr, uint64_t src_vaddr, uint64_t si
}
auto& command = m_scheduler.Current();
auto src = src_memory ? ObtainBuffer(command, src_vaddr, size, false, true)
auto src = src_memory ? ObtainBuffer(command, src_vaddr, size, false, true, true)
: BufferBinding {.buffer = m_gds_buffer.Handle(), .offset = src_vaddr};
auto dst = dst_memory ? ObtainBuffer(command, dst_vaddr, size, true, false, true)
: BufferBinding {.buffer = m_gds_buffer.Handle(), .offset = dst_vaddr};
@@ -1134,95 +837,13 @@ bool BufferCache::IsRegionCpuModified(uint64_t vaddr, uint64_t size) {
return m_memory_tracker.IsRegionCpuModified(vaddr, size);
}
void BufferCache::InvalidateImageAliases(uint64_t vaddr, uint64_t size) {
if (vaddr == 0 || size == 0 || vaddr >= TRACKER_ADDRESS_SIZE ||
size > TRACKER_ADDRESS_SIZE - vaddr) {
EXIT("BufferCache: invalid image-alias invalidation\n");
}
FaultSafeCacheLock lock(this, m_mutex);
const auto end = vaddr + size;
for (const auto& [address, cached]: m_buffers) {
const auto cached_end = address + cached->size;
const auto begin = std::max(vaddr, address);
const auto range_end = std::min(end, cached_end);
if (begin >= range_end) {
continue;
}
const auto bytes = range_end - begin;
if (!m_gpu_modified_ranges.Intersections(begin, bytes).empty()) {
EXIT("BufferCache: image ownership overlaps exact dirty buffer bytes\n");
}
m_image_invalidated_ranges.Add(begin, bytes);
}
}
void BufferCache::BeginBackingPublication(uint64_t vaddr, uint64_t size, uint64_t tick) {
if (vaddr == 0 || size == 0 || tick == 0 || vaddr >= TRACKER_ADDRESS_SIZE ||
size > TRACKER_ADDRESS_SIZE - vaddr) {
EXIT("BufferCache: invalid pending backing publication\n");
}
std::lock_guard lock(m_publication_mutex);
m_pending_backing_publications.push_back({vaddr, size, tick});
}
void BufferCache::CompleteBackingPublication(uint64_t vaddr, uint64_t size, uint64_t tick) {
std::lock_guard lock(m_publication_mutex);
const auto publication =
std::ranges::find_if(m_pending_backing_publications, [&](const auto& pending) {
return pending.address == vaddr && pending.size == size && pending.tick == tick;
});
if (publication == m_pending_backing_publications.end()) {
EXIT("BufferCache: completed an unknown backing publication\n");
}
m_pending_backing_publications.erase(publication);
}
void BufferCache::PublishImageBuffer(uint64_t vaddr, uint64_t size) {
FaultSafeCacheLock lock(this, m_mutex);
auto owner = m_buffers.end();
for (auto it = m_buffers.begin(); it != m_buffers.end(); ++it) {
if (!PageOverlaps(vaddr, size, it->second->vaddr, it->second->size)) {
continue;
}
if (owner != m_buffers.end() || !it->second->buffer->IsInBounds(vaddr, size)) {
EXIT("BufferCache: image destination aliases a non-containing cached buffer\n");
}
owner = it;
}
m_memory_tracker.ValidateGpuDirtyOwnership(m_gpu_modified_ranges, vaddr, size,
"image destination publication");
if (owner == m_buffers.end() || m_memory_tracker.IsRegionCpuModified(vaddr, size) ||
!m_gpu_modified_ranges.Intersections(vaddr, size).empty()) {
EXIT("BufferCache: image destination requires clean buffer ownership\n");
}
m_memory_tracker.MarkRegionAsGpuModified(vaddr, size);
m_gpu_modified_ranges.Add(vaddr, size);
m_image_invalidated_ranges.Subtract(vaddr, size);
m_memory_tracker.ValidateGpuDirtyOwnership(m_gpu_modified_ranges, vaddr, size,
"published image destination");
owner->second->tick_accessed_last = m_gc_tick;
}
void BufferCache::ValidateGpuAccess(uint64_t vaddr, uint64_t size, bool is_read,
bool is_written) const {
if ((!is_read && !is_written) || vaddr == 0 || size == 0 || size > UINT64_MAX - vaddr) {
EXIT("BufferCache: invalid GPU access request\n");
}
if (is_read && !m_page_manager.HasGpuAccess(vaddr, size, GpuAccess::Read)) {
EXIT("BufferCache: GPU-read access denied\n");
}
if (is_written && !m_page_manager.HasGpuAccess(vaddr, size, GpuAccess::Write)) {
EXIT("BufferCache: GPU-write access denied\n");
}
}
void BufferCache::RunGarbageCollector() {
std::lock_guard transaction(m_resource_mutex);
const auto tick = m_gc_tick++;
if (m_graphics.CanReportMemoryUsage()) {
m_total_used_memory = m_graphics.GetDeviceMemoryUsage();
}
if (m_total_used_memory < m_trigger_gc_memory || m_fault_readback->Active()) {
if (m_total_used_memory < m_trigger_gc_memory) {
return;
}
@@ -1291,7 +912,6 @@ void BufferCache::RunGarbageCollector() {
if (!m_memory_tracker.IsRegionGpuModified(retire.address, retire.size)) {
m_memory_tracker.UntrackMemory(retire.address, retire.size);
}
m_image_invalidated_ranges.Subtract(retire.address, retire.size);
if (retire.size > m_total_used_memory) {
EXIT("BufferCache: allocation accounting underflow\n");
}
+7 -29
View File
@@ -10,7 +10,6 @@
#include <map>
#include <memory>
#include <mutex>
#include <span>
#include <utility>
#include <vector>
@@ -30,9 +29,8 @@ struct BufferBinding {
};
struct ImageBufferSource {
Buffer* buffer = nullptr;
uint64_t offset = 0;
bool gpu_owned = false;
Buffer* buffer = nullptr;
uint64_t offset = 0;
};
class BufferCache {
@@ -47,11 +45,9 @@ public:
~BufferCache();
KYTY_CLASS_NO_COPY(BufferCache);
[[nodiscard]] bool InvalidateMemory(PageFaultAccess access, uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept;
void InvalidateMemory(uint64_t vaddr, uint64_t size);
void ReadMemory(uint64_t vaddr, uint64_t size);
void UnmapMemory(uint64_t vaddr, uint64_t size);
void InvalidateMemory(uint64_t vaddr, uint64_t size);
void ReadMemory(uint64_t vaddr, uint64_t size);
void UnmapMemory(uint64_t vaddr, uint64_t size);
[[nodiscard]] BufferBinding ObtainBuffer(CommandBuffer& command, uint64_t vaddr, uint64_t size,
bool is_written = false, bool is_read = true,
bool is_formatted = false);
@@ -62,9 +58,6 @@ public:
uint64_t alignment);
[[nodiscard]] std::shared_ptr<Buffer> ObtainNullBuffer();
[[nodiscard]] ImageBufferSource ObtainBufferForImage(uint64_t vaddr, uint64_t size);
[[nodiscard]] std::pair<std::shared_ptr<Buffer>, uint64_t>
ObtainBufferForImageWrite(uint64_t vaddr, uint64_t size);
void DiscardGpuDirtyBytes(uint64_t vaddr, uint64_t size);
void FillBuffer(uint64_t vaddr, uint64_t size, uint32_t value, bool is_gds = false);
void CopyBuffer(uint64_t dst_vaddr, uint64_t src_vaddr, uint64_t size, bool dst_gds = false,
bool src_gds = false);
@@ -72,13 +65,7 @@ public:
[[nodiscard]] bool HasGpuDirtyBytes(uint64_t vaddr, uint64_t size);
[[nodiscard]] bool IsRegionCpuModified(uint64_t vaddr, uint64_t size);
[[nodiscard]] bool IsRegionGpuModified(uint64_t vaddr, uint64_t size);
void InvalidateImageAliases(uint64_t vaddr, uint64_t size);
void BeginBackingPublication(uint64_t vaddr, uint64_t size, uint64_t tick);
void CompleteBackingPublication(uint64_t vaddr, uint64_t size, uint64_t tick);
[[nodiscard]] bool SynchronizeBacking(uint64_t vaddr, uint64_t size);
void PublishImageBuffer(uint64_t vaddr, uint64_t size);
void ValidateGpuAccess(uint64_t vaddr, uint64_t size, bool is_read, bool is_written) const;
void RunGarbageCollector();
void RunGarbageCollector();
private:
friend struct BufferCacheTestAccess;
@@ -91,8 +78,6 @@ private:
struct DownloadCopy;
struct DownloadRange;
struct RetiredBuffer;
struct FaultReadback;
struct PendingBackingPublication;
static constexpr uint64_t DOWNLOAD_ALIGNMENT = 64;
[[nodiscard]] static uint64_t AlignDown(uint64_t value) noexcept;
[[nodiscard]] static uint64_t AlignUp(uint64_t value);
@@ -107,12 +92,10 @@ private:
const void* source, uint64_t size);
[[nodiscard]] CachedBuffer& GetOrCreateBuffer(CommandBuffer& command, uint64_t vaddr,
uint64_t size);
[[nodiscard]] bool SynchronizeBufferFromImage(Buffer& buffer, uint64_t vaddr, uint64_t size);
[[nodiscard]] std::vector<DownloadRange> RecordDownloads(std::span<const DownloadCopy> copies);
void PublishDownloads(std::span<const DownloadRange> downloads);
void QueueGarbageDownload(std::span<const DownloadCopy> copies, RetiredBuffer retire);
void RefreshInvalidatedRanges(CommandBuffer& command, CachedBuffer& cached, uint64_t vaddr,
uint64_t size, bool upload);
void DiscardGpuDirtyBytesLocked(uint64_t vaddr, uint64_t size, const char* operation);
void WriteHostMemory(uint64_t vaddr, std::span<const uint8_t> data);
GraphicContext& m_graphics;
@@ -121,17 +104,12 @@ private:
Common::Mutex m_mutex;
std::shared_ptr<Buffer> m_null_buffer;
std::map<uint64_t, std::unique_ptr<CachedBuffer>> m_buffers;
std::unique_ptr<FaultReadback> m_fault_readback;
RangeSet m_gpu_modified_ranges;
RangeSet m_image_invalidated_ranges;
std::mutex m_publication_mutex;
std::vector<PendingBackingPublication> m_pending_backing_publications;
MemoryTracker m_memory_tracker;
StreamBuffer m_staging_buffer;
StreamBuffer m_stream_buffer;
StreamBuffer m_download_buffer;
StreamBuffer m_device_buffer;
PageManager& m_page_manager;
TextureCache& m_texture_cache;
ResourceMutex& m_resource_mutex;
uint64_t m_total_used_memory = 0;
+21 -35
View File
@@ -4,37 +4,15 @@
#include "graphics/guest_gpu/command_processor/commandProcessor.h"
#include "graphics/guest_gpu/graphicsRun.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
namespace Libs::Graphics {
GpuResourceManager::GpuResourceManager(GraphicContext& graphics, CommandScheduler& scheduler)
: m_page_manager(FaultThunk, this),
: m_scheduler(scheduler),
m_buffer_cache(graphics, scheduler, m_page_manager, m_texture_cache, m_resource_mutex),
m_texture_cache(graphics, scheduler, m_page_manager, m_buffer_cache, m_resource_mutex) {}
GpuResourceManager::~GpuResourceManager() = default;
bool GpuResourceManager::FaultThunk(void* context, PageFaultAccess access, uint64_t vaddr,
uint64_t size, PageFaultPhase phase) noexcept {
return static_cast<GpuResourceManager*>(context)->InvalidateMemory(access, vaddr, size, phase);
}
bool GpuResourceManager::InvalidateMemory(PageFaultAccess access, uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept {
// Let the authoritative image materialize first. A clean overlapping buffer marks a write
// fault CPU-dirty when it begins ownership transfer; doing that before image preflight would
// make the image appear to race a real CPU write. Completion and release retain buffer-first
// ordering so its pending fault is gone before TextureCache publishes the downloaded backing.
if (phase == PageFaultPhase::Invalidate) {
const bool image_handled = m_texture_cache.InvalidateMemory(access, vaddr, size, phase);
const bool buffer_handled = m_buffer_cache.InvalidateMemory(access, vaddr, size, phase);
return buffer_handled || image_handled;
}
const bool buffer_handled = m_buffer_cache.InvalidateMemory(access, vaddr, size, phase);
const bool image_handled = m_texture_cache.InvalidateMemory(access, vaddr, size, phase);
return buffer_handled || image_handled;
}
bool GpuResourceManager::HandleFault(PageFaultAccess access, uint64_t fault_vaddr) noexcept {
constexpr uint64_t fault_size = 8;
if (!IsMapped(fault_vaddr, fault_size)) {
@@ -115,33 +93,41 @@ bool GpuResourceManager::IsMapped(uint64_t vaddr, uint64_t size) const noexcept
return m_mapped_ranges.Contains(vaddr, size);
}
void GpuResourceManager::MapMemory(uint64_t vaddr, uint64_t size, GpuAccess access) {
void GpuResourceManager::MapMemory(uint64_t vaddr, uint64_t size) {
{
std::lock_guard lock(m_mapped_ranges_mutex);
m_mapped_ranges.Add(vaddr, size);
}
m_page_manager.OnGpuMap(vaddr, size, access);
m_page_manager.OnGpuMap(vaddr, size);
}
void GpuResourceManager::UnmapMemory(uint64_t vaddr, uint64_t size, GpuAccess access) {
if (!IsMapped(vaddr, size)) {
EXIT("cannot unmap an unmapped GPU resource range\n");
void GpuResourceManager::UnmapMemory(uint64_t vaddr, uint64_t size) {
if (CommandScheduler::InDeferredOperation()) {
EXIT("unsupported memory unmap from an asynchronous GPU completion, "
"addr=0x%016" PRIx64 " size=0x%016" PRIx64 "\n",
vaddr, size);
}
const auto unmap = [this, vaddr, size, access] {
m_texture_cache.UnmapMemory(vaddr, size);
if (m_resource_mutex.IsOwnedByCurrentThread()) {
EXIT("unsupported memory unmap from a pre-owned resource transaction, "
"addr=0x%016" PRIx64 " size=0x%016" PRIx64 "\n",
vaddr, size);
}
const auto unmap = [this, vaddr, size] {
if (m_scheduler.Active()) {
const auto tick = m_scheduler.CurrentTick();
m_scheduler.FinishCurrent();
m_scheduler.WaitPriorityOperations(tick);
}
m_buffer_cache.UnmapMemory(vaddr, size);
m_page_manager.OnGpuUnmap(vaddr, size, access);
m_texture_cache.UnmapMemory(vaddr, size);
m_page_manager.OnGpuUnmap(vaddr, size);
std::lock_guard lock(m_mapped_ranges_mutex);
m_mapped_ranges.Subtract(vaddr, size);
};
if (m_gpu == nullptr) {
if (m_resource_mutex.IsOwnedByCurrentThread()) {
EXIT("cannot synchronously unmap from a resource transaction\n");
}
unmap();
return;
}
Gpu::SubmissionLock submissions(*m_gpu);
m_gpu->SendCommandSync(unmap);
}
+3 -7
View File
@@ -29,18 +29,14 @@ public:
[[nodiscard]] bool HandleFault(PageFaultAccess access, uint64_t fault_vaddr) noexcept;
[[nodiscard]] bool InvalidateMemory(uint64_t vaddr, uint64_t size);
[[nodiscard]] bool IsMapped(uint64_t vaddr, uint64_t size) const noexcept;
void MapMemory(uint64_t vaddr, uint64_t size, GpuAccess access);
void UnmapMemory(uint64_t vaddr, uint64_t size, GpuAccess access);
void MapMemory(uint64_t vaddr, uint64_t size);
void UnmapMemory(uint64_t vaddr, uint64_t size);
void RunGarbageCollector();
private:
static bool FaultThunk(void* context, PageFaultAccess access, uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept;
[[nodiscard]] bool InvalidateMemory(PageFaultAccess access, uint64_t vaddr, uint64_t size,
PageFaultPhase phase) noexcept;
PageManager m_page_manager;
ResourceMutex m_resource_mutex;
CommandScheduler& m_scheduler;
BufferCache m_buffer_cache;
TextureCache m_texture_cache;
mutable std::shared_mutex m_mapped_ranges_mutex;
+3 -3
View File
@@ -219,7 +219,7 @@ private:
typename CoarseTable::PageRange coarse_range {};
typename TrackingTable::PageRange tracking_range {};
if (!CoarseTable::TryGetPageRange(address, size, coarse_range) ||
!TrackingTable::TryGetPageRange(address, size, tracking_range)) {
(!strict_bytes && !TrackingTable::TryGetPageRange(address, size, tracking_range))) {
return {};
}
MembershipList candidates;
@@ -230,8 +230,8 @@ private:
}
std::vector<OwnerT> result;
for (const Registration* registration: candidates) {
if ((!strict_bytes || Overlaps(registration->ranges, address, size)) &&
HasTrackingMembership(registration, tracking_range) &&
if ((strict_bytes ? Overlaps(registration->ranges, address, size)
: HasTrackingMembership(registration, tracking_range)) &&
predicate(registration->owner)) {
result.push_back(registration->owner);
}
+4 -1
View File
@@ -56,7 +56,10 @@ vk::Sampler SamplerCache::GetSampler(const ShaderSamplerResource& r) {
case Prospero::SamplerAnisoRatio::kFour: aniso_ratio = 4.0f; break;
case Prospero::SamplerAnisoRatio::kEight: aniso_ratio = 8.0f; break;
case Prospero::SamplerAnisoRatio::kSixteen: aniso_ratio = 16.0f; break;
default: EXIT("unknown ratio: %d\n", static_cast<int>(r.MaxAnisoRatio()));
default:
EXIT("unknown ratio: %d dwords=%08x,%08x,%08x,%08x\n",
static_cast<int>(r.MaxAnisoRatio()), r.fields[0], r.fields[1], r.fields[2],
r.fields[3]);
}
}
+10 -12
View File
@@ -135,8 +135,8 @@ void Buffer::Write(uint64_t offset, const void* source, uint64_t size) {
void Buffer::Flush(uint64_t offset, uint64_t size) {
EXIT_IF(m_mapped.empty() || offset > m_size || size > m_size - offset);
if (!m_is_coherent && size != 0) {
const auto result = vmaFlushAllocation(m_graphics->allocator, m_buffer->memory.allocation,
offset, size);
const auto result =
vmaFlushAllocation(m_graphics->allocator, m_buffer->memory.allocation, offset, size);
EXIT_NOT_IMPLEMENTED(static_cast<vk::Result>(result) != vk::Result::eSuccess);
}
}
@@ -144,8 +144,8 @@ void Buffer::Flush(uint64_t offset, uint64_t size) {
vk::BufferMemoryBarrier Buffer::Barrier(uint64_t offset, uint64_t size, vk::AccessFlags source,
vk::AccessFlags destination) const {
if (Handle() == nullptr || size == 0 || offset > m_size || size > m_size - offset) {
EXIT("Buffer: invalid DMA barrier, handle=%p offset=0x%016" PRIx64
" size=0x%016" PRIx64 " capacity=0x%016" PRIx64 "\n",
EXIT("Buffer: invalid DMA barrier, handle=%p offset=0x%016" PRIx64 " size=0x%016" PRIx64
" capacity=0x%016" PRIx64 "\n",
static_cast<const void*>(Handle()), offset, size, m_size);
}
vk::BufferMemoryBarrier barrier {};
@@ -175,10 +175,9 @@ void Buffer::CopyFrom(CommandBuffer& command, const Buffer& source, uint64_t sou
command.EndRendering();
const vk::BufferMemoryBarrier before[] = {
source.Barrier(source_offset, size, source_before, vk::AccessFlagBits::eTransferRead),
Barrier(destination_offset, size, destination_before,
vk::AccessFlagBits::eTransferWrite),
Barrier(destination_offset, size, destination_before, vk::AccessFlagBits::eTransferWrite),
};
const auto host_access = vk::AccessFlagBits::eHostRead | vk::AccessFlagBits::eHostWrite;
const auto host_access = vk::AccessFlagBits::eHostRead | vk::AccessFlagBits::eHostWrite;
auto before_stage = vk::PipelineStageFlags {vk::PipelineStageFlagBits::eAllCommands};
if (static_cast<bool>((source_before | destination_before) & host_access)) {
before_stage |= vk::PipelineStageFlagBits::eHost;
@@ -214,9 +213,8 @@ void Buffer::Fill(uint64_t offset, uint64_t size, uint32_t value) {
vk::PipelineStageFlagBits::eTransfer, vk::DependencyFlagBits::eByRegion,
0, nullptr, 1, &before, 0, nullptr);
native.fillBuffer(Handle(), offset, size, value);
const auto after =
Barrier(offset, size, vk::AccessFlagBits::eTransferWrite,
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite);
const auto after = Barrier(offset, size, vk::AccessFlagBits::eTransferWrite,
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite);
native.pipelineBarrier(vk::PipelineStageFlagBits::eTransfer,
vk::PipelineStageFlagBits::eAllCommands,
vk::DependencyFlagBits::eByRegion, 0, nullptr, 1, &after, 0, nullptr);
@@ -250,8 +248,8 @@ std::pair<uint8_t*, uint64_t> StreamBuffer::Map(uint64_t size, uint64_t alignmen
if (Mapped().empty()) {
return {nullptr, 0};
}
uint64_t mapped_size = size;
const auto atom = Graphics().physical_device_properties.limits.nonCoherentAtomSize;
uint64_t mapped_size = size;
const auto atom = Graphics().physical_device_properties.limits.nonCoherentAtomSize;
if (!NormalizeReservation(IsCoherent(), atom, mapped_size, alignment)) {
return {nullptr, 0};
}
+14 -15
View File
@@ -54,16 +54,15 @@ public:
[[nodiscard]] bool IsInBounds(uint64_t address, uint64_t size) const noexcept;
void Write(uint64_t offset, const void* source, uint64_t size);
void Flush(uint64_t offset, uint64_t size);
void CopyFrom(
CommandBuffer& command, const Buffer& source, uint64_t source_offset,
uint64_t destination_offset, uint64_t size,
vk::AccessFlags source_before = vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_before =
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags source_after =
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_after =
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite);
void CopyFrom(CommandBuffer& command, const Buffer& source, uint64_t source_offset,
uint64_t destination_offset, uint64_t size,
vk::AccessFlags source_before = vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_before = vk::AccessFlagBits::eMemoryRead |
vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags source_after = vk::AccessFlagBits::eMemoryRead |
vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_after = vk::AccessFlagBits::eMemoryRead |
vk::AccessFlagBits::eMemoryWrite);
void Fill(uint64_t offset, uint64_t size, uint32_t value);
protected:
@@ -107,13 +106,13 @@ private:
uint64_t upper_bound = 0;
};
void ReserveWatches(std::vector<Watch>& watches, size_t grow_size);
void ReserveWatches(std::vector<Watch>& watches, size_t grow_size);
[[nodiscard]] static bool NormalizeReservation(bool coherent, uint64_t atom, uint64_t& size,
uint64_t& alignment);
[[nodiscard]] bool WaitPendingOperations(const std::vector<Watch>& watches,
std::optional<size_t> invalidation_mark,
uint64_t requested_upper_bound, bool allow_wait,
size_t& wait_cursor, uint64_t& wait_bound);
[[nodiscard]] bool WaitPendingOperations(const std::vector<Watch>& watches,
std::optional<size_t> invalidation_mark,
uint64_t requested_upper_bound, bool allow_wait,
size_t& wait_cursor, uint64_t& wait_bound);
uint64_t m_offset = 0;
uint64_t m_mapped_size = 0;
+133 -141
View File
@@ -323,7 +323,7 @@ void TextureCache::TrackImage(ImageId id) {
if (!image.IsTracked()) {
image.track_addr = image_begin;
image.track_addr_end = image_end;
m_page_manager.UpdatePageWatchers(true, image_begin, image.info.data.size);
m_page_manager.UpdatePageWatchers<true>(image_begin, image.info.data.size);
return;
}
if (image_begin < image.track_addr) {
@@ -348,7 +348,7 @@ void TextureCache::TrackImageHead(ImageId id) {
}
const auto size = image.track_addr - image_begin;
image.track_addr = image_begin;
m_page_manager.UpdatePageWatchers(true, image_begin, size);
m_page_manager.UpdatePageWatchers<true>(image_begin, size);
}
void TextureCache::TrackImageTail(ImageId id) {
@@ -366,7 +366,7 @@ void TextureCache::TrackImageTail(ImageId id) {
const auto address = image.track_addr_end;
const auto size = image_end - address;
image.track_addr_end = image_end;
m_page_manager.UpdatePageWatchers(true, address, size);
m_page_manager.UpdatePageWatchers<true>(address, size);
}
void TextureCache::UntrackImage(ImageId id) {
@@ -379,7 +379,7 @@ void TextureCache::UntrackImage(ImageId id) {
image.track_addr = 0;
image.track_addr_end = 0;
if (size != 0) {
m_page_manager.UpdatePageWatchers(false, address, size);
m_page_manager.UpdatePageWatchers<false>(address, size);
}
}
@@ -400,7 +400,7 @@ void TextureCache::UntrackImageHead(ImageId id) {
UntrackImage(id);
}
if (size != 0) {
m_page_manager.UpdatePageWatchers(false, begin, size);
m_page_manager.UpdatePageWatchers<false>(begin, size);
}
}
@@ -421,7 +421,7 @@ void TextureCache::UntrackImageTail(ImageId id) {
UntrackImage(id);
}
if (size != 0) {
m_page_manager.UpdatePageWatchers(false, address, size);
m_page_manager.UpdatePageWatchers<false>(address, size);
}
}
@@ -867,22 +867,28 @@ TextureCache::BuildColorTransfer(const Image& image, BindingType binding,
case BindingType::Texture: break;
case BindingType::Storage: owner = "StorageTextureCache"; break;
case BindingType::RenderTarget:
if (info.resources.layers == 0 || info.data.size % info.resources.layers != 0 ||
info.samples != 1 || image.backing.samples != 1) {
EXIT("TextureCache: invalid color-attachment upload\n");
}
format = ImageOps::RenderTargetTransferFormat(info.bytes_per_block);
allow_depth_tile = false;
plan.swap_bgra16 = info.bgra16;
owner = "RenderTarget";
break;
case BindingType::VideoOut:
if (info.resources.layers == 0 || info.data.size % info.resources.layers != 0 ||
info.samples != 1 || image.backing.samples != 1 ||
(binding == BindingType::VideoOut &&
info.metadata.compression != VideoOutCompression::Uncompressed)) {
info.metadata.compression != VideoOutCompression::Uncompressed) {
EXIT("TextureCache: invalid color-attachment upload\n");
}
format = binding == BindingType::RenderTarget
? ImageOps::RenderTargetTransferFormat(info.bytes_per_block)
: info.guest_format;
format = info.guest_format;
layers = info.resources.layers;
volume = false;
layered = layers > 1;
allow_depth_tile = false;
plan.swap_bgra16 = info.bgra16;
owner = binding == BindingType::RenderTarget ? "RenderTarget" : "VideoOut";
owner = "VideoOut";
break;
case BindingType::DepthTarget: return plan;
}
@@ -1038,25 +1044,20 @@ void TextureCache::InitializeImage(ImageId id, const ImageDesc& desc) {
if (image.info.samples > 1) {
return;
}
bool data_gpu_owned = false;
bool data_imported = false;
const bool upload = image.IsBufferModified() || image.IsCpuDirty();
bool data_imported = false;
const bool upload = image.IsBufferModified() || image.IsCpuDirty();
if (upload) {
const auto source =
m_buffer_cache.ObtainBufferForImage(image.info.data.address, image.info.data.size);
if (source.buffer == nullptr) {
EXIT("TextureCache: failed to obtain image upload source\n");
}
data_gpu_owned |= source.gpu_owned;
data_imported = true;
UploadImage(image, desc, *source.buffer, source.offset);
}
if (data_imported) {
image.ClearBufferModified();
}
if (data_gpu_owned) {
image.MarkGpuModified();
}
if (image.IsCpuDirty()) {
image.RefreshComplete();
}
@@ -1131,8 +1132,7 @@ ImageId TextureCache::FindImage(ImageDesc& desc, bool exact_format) {
}
ImageId result {};
bool replacement_buffer = false;
bool inserted_new = false;
bool inserted_new = false;
{
std::lock_guard transaction(m_resource_mutex);
CacheLock lock(*this, m_lock);
@@ -1159,8 +1159,8 @@ ImageId TextureCache::FindImage(ImageDesc& desc, bool exact_format) {
if (owner == nullptr) {
continue;
}
const auto merged_info = result ? ResolveImage(result).info : desc.info;
const auto overlap = ResolveOverlap(merged_info, desc.type, candidate, result);
const auto& merged_info = result ? ResolveImage(result).info : desc.info;
const auto overlap = ResolveOverlap(merged_info, desc.type, candidate, result);
if (overlap.image) {
result = overlap.image;
view_mip = overlap.mip;
@@ -1174,23 +1174,15 @@ ImageId TextureCache::FindImage(ImageDesc& desc, bool exact_format) {
if (exact_format && resolved.info.pixel_format != desc.info.pixel_format) {
result = {};
} else if (resolved.info.resources < desc.info.resources) {
ImageDesc refresh {
.info = resolved.info, .view_info = {}, .type = UploadBinding(resolved)};
RefreshImage(result, refresh);
if (resolved.IsGpuModified() && !SynchronizeImageToBuffer(result)) {
EXIT("TextureCache: cannot preserve an unsupported replacement image\n");
}
replacement_buffer = resolved.IsBufferModified();
DeleteImage(result);
result = {};
result = ExpandImage(desc.info, result);
}
}
if (!result) {
result = InsertImage(desc.info);
inserted_new = true;
auto& inserted = ResolveImage(result);
if (replacement_buffer || m_buffer_cache.HasGpuDirtyBytes(inserted.info.data.address,
inserted.info.data.size)) {
if (m_buffer_cache.HasGpuDirtyBytes(inserted.info.data.address,
inserted.info.data.size)) {
inserted.MarkBufferModified();
}
}
@@ -1360,11 +1352,6 @@ void TextureCache::CommitGpuWrite(Image& image) {
if (image.depth_id || image.backing.image == nullptr) {
EXIT("TextureCache: stencil association cannot own image contents\n");
}
const auto range = image.info.data;
if (m_buffer_cache.HasGpuDirtyBytes(range.address, range.size)) {
m_buffer_cache.DiscardGpuDirtyBytes(range.address, range.size);
}
m_buffer_cache.InvalidateImageAliases(range.address, range.size);
image.ClearBufferModified();
if (image.IsCpuDirty()) {
image.RefreshComplete();
@@ -1377,7 +1364,6 @@ bool TextureCache::ClearImageFromBuffer(CommandBuffer& command, uint64_t address
if (command.IsInvalid() || !GuestRange {address, size}.Valid()) {
EXIT("TextureCache: invalid image clear\n");
}
m_buffer_cache.ValidateGpuAccess(address, size, false, true);
std::lock_guard transaction(m_resource_mutex);
CacheLock lock(*this, m_lock);
ImageId selected {};
@@ -1430,9 +1416,6 @@ bool TextureCache::ClearImageFromBuffer(CommandBuffer& command, uint64_t address
return false;
}
}
if (m_buffer_cache.HasGpuDirtyBytes(address, size)) {
m_buffer_cache.DiscardGpuDirtyBytes(address, size);
}
if (image.IsBufferModified() || image.IsCpuDirty()) {
ImageDesc refresh {.info = image.info, .view_info = {}, .type = UploadBinding(image)};
InitializeImage(selected, refresh);
@@ -1552,11 +1535,14 @@ void TextureCache::DownloadDepth(Image& image, Buffer& destination, uint64_t des
}
void TextureCache::DownloadImageData(Image& image, Buffer& destination, uint64_t destination_offset,
DownloadPlan plan) {
uint64_t destination_size, DownloadPlan plan) {
if (!plan.valid) {
EXIT("TextureCache: invalid image download plan\n");
}
if (plan.depth) {
if (destination_size != image.info.data.size) {
EXIT("TextureCache: partial depth image download is unsupported\n");
}
DownloadDepth(image, destination, destination_offset);
return;
}
@@ -1566,22 +1552,118 @@ void TextureCache::DownloadImageData(Image& image, Buffer& destination, uint64_t
: TileManager::ColorTransform::None;
if (!color.tiled) {
if (transform == TileManager::ColorTransform::SwapBgra16) {
auto linear = m_tiler->GetScratchBuffer(image.info.data.size);
auto linear = m_tiler->GetScratchBuffer(destination_size);
image.Download(color.regions, linear.buffer, 0, linear.size);
m_tiler->SwapBgra16(linear,
{destination.Handle(), destination_offset, image.info.data.size});
{destination.Handle(), destination_offset, destination_size});
return;
}
for (auto& copy: color.regions) {
copy.bufferOffset += destination_offset;
}
image.Download(color.regions, destination.Handle(), destination_offset,
image.info.data.size);
image.Download(color.regions, destination.Handle(), destination_offset, destination_size);
return;
}
m_tiler->TileImage(image, color.regions, destination.Handle(), destination_offset,
image.info.data.size, image.info.data.size, color.tiles, transform);
destination_size, destination_size, color.tiles, transform);
}
bool BufferCache::SynchronizeBufferFromImage(Buffer& buffer, uint64_t vaddr, uint64_t size) {
CacheLock lock(m_texture_cache, m_texture_cache.m_lock);
std::vector<ImageId> matches;
for (const auto id: m_texture_cache.FindImagesInRegion(vaddr, size, false)) {
auto owner = m_texture_cache.ResolveOwner(id);
if (owner == nullptr || owner->info.data.address != vaddr) {
continue;
}
if (owner->depth_id) {
owner = m_texture_cache.ResolveOwner(owner->depth_id);
}
if (owner != nullptr && owner->SafeToDownload()) {
matches.push_back(id);
}
}
ImageId selected {};
if (matches.size() == 1) {
selected = matches.front();
} else {
for (const auto id: matches) {
const auto& image = m_texture_cache.ResolveImage(id);
if (image.info.data.size == size) {
selected = id;
break;
}
}
}
if (!selected) {
return false;
}
if (const auto owner = m_texture_cache.ResolveOwner(selected);
owner != nullptr && owner->depth_id) {
selected = owner->depth_id;
}
auto& image = m_texture_cache.ResolveImage(selected);
if (!buffer.IsInBounds(image.info.data.address, 1)) {
return false;
}
const auto buf_offset = buffer.Offset(image.info.data.address);
const auto available = buffer.Size() - buf_offset;
uint32_t levels = 0;
uint64_t copy_size = 0;
if (image.info.IsVolume()) {
// Volume mips contain strided block slices, so a mip's linear span cannot prove that
// every retained slice fits. Keep volume synchronization whole-image only.
if (!buffer.IsInBounds(image.info.data.address, image.info.data.size)) {
return false;
}
levels = image.info.resources.levels;
copy_size = image.info.data.size;
} else {
for (; levels < image.info.resources.levels; ++levels) {
const auto& mip = image.info.mip_layout[levels];
if (mip.size == 0 || mip.offset > available || mip.size > available - mip.offset) {
break;
}
copy_size = std::max(copy_size, mip.offset + mip.size);
}
}
if (copy_size == 0) {
return false;
}
auto plan = m_texture_cache.BuildDownload(image);
if (!plan.valid) {
return false;
}
if (plan.depth && copy_size != image.info.data.size) {
return false;
}
if (!plan.depth && levels < image.info.resources.levels) {
auto& color = plan.color;
std::erase_if(color.regions, [levels](const vk::BufferImageCopy& region) {
return region.imageSubresource.mipLevel >= levels;
});
if (color.regions.empty()) {
return false;
}
if (color.tiled) {
const auto binding = m_texture_cache.UploadBinding(image);
const auto format =
binding == TextureCache::BindingType::RenderTarget
? ImageOps::RenderTargetTransferFormat(image.info.bytes_per_block)
: image.info.guest_format;
color.tiles.clear();
if (!TextureBuildGpuTileInfos(copy_size, color.regions, color.layout, format,
image.info.TransferLayers(), levels, color.tiles)) {
return false;
}
}
}
m_texture_cache.DownloadImageData(image, buffer, buf_offset, copy_size, std::move(plan));
m_texture_cache.RetainImage(m_scheduler.Current(), selected);
return true;
}
std::pair<uint8_t*, uint64_t> TextureCache::MapDownload(uint64_t size, uint64_t alignment) {
@@ -1613,12 +1695,9 @@ void TextureCache::QueueDownload(GuestRange range, StreamBuffer& download, uint8
m_scheduler.Current().Handle().pipelineBarrier(vk::PipelineStageFlagBits::eAllCommands,
vk::PipelineStageFlagBits::eHost, {}, 0, nullptr,
1, &barrier, 0, nullptr);
const auto tick = m_scheduler.CurrentTick();
m_buffer_cache.BeginBackingPublication(range.address, range.size, tick);
m_scheduler.DeferPriorityOperation([this, &download, range, mapped, offset, tick] {
m_scheduler.DeferPriorityOperation([&download, range, mapped, offset] {
download.Invalidate(offset, range.size);
LibKernel::Memory::WriteBacking(range.address, mapped, range.size);
m_buffer_cache.CompleteBackingPublication(range.address, range.size, tick);
});
}
@@ -1639,7 +1718,7 @@ bool TextureCache::TryDownloadImage(ImageId id) {
}
download.Flush(offset, range.size);
DownloadImageData(image, download, offset, std::move(plan));
DownloadImageData(image, download, offset, range.size, std::move(plan));
QueueDownload(range, download, mapped, offset);
return true;
@@ -1653,62 +1732,6 @@ void TextureCache::DownloadImage(ImageId id) {
m_scheduler.DrainPriorityOperations();
}
bool TextureCache::SynchronizeImageToBuffer(ImageId id) {
auto& image = ResolveImage(id);
if (image.depth_id) {
return true;
}
auto plan = BuildDownload(image);
if (!plan.valid) {
return false;
}
const auto range = image.info.data;
if (image.IsCpuDirty()) {
RefreshImage(id,
ImageDesc {.info = image.info, .view_info = {}, .type = UploadBinding(image)});
}
if (!image.IsGpuModified()) {
return true;
}
if (image.IsDefinitelyCpuDirty() || image.IsBufferModified()) {
EXIT("TextureCache: image mirror source is not native-current\n");
}
auto [destination, offset] =
m_buffer_cache.ObtainBufferForImageWrite(range.address, range.size);
if (destination == nullptr) {
EXIT("TextureCache: failed to allocate image mirror\n");
}
DownloadImageData(image, *destination, offset, std::move(plan));
m_scheduler.Current().RetainResourceUntilFence(destination);
m_buffer_cache.PublishImageBuffer(range.address, range.size);
image.MarkBufferModified();
RetainImage(m_scheduler.Current(), id);
ClearGpuModified(id);
return true;
}
bool TextureCache::SynchronizeImageToBuffer(uint64_t address, uint64_t size) {
if (!GuestRange {address, size}.Valid()) {
return false;
}
CacheLock lock(*this, m_lock);
ImageId selected {};
for (const auto id: FindImagesInRegion(address, size, true)) {
auto owner = ResolveOwner(id);
if (owner == nullptr || !owner->GpuOverlaps(address, size)) {
continue;
}
if (selected) {
EXIT("TextureCache: ambiguous image-to-buffer synchronization\n");
}
selected = id;
}
if (!selected) {
return false;
}
return SynchronizeImageToBuffer(selected);
}
bool TextureCache::InvalidateMemoryFromGPU(uint64_t address, uint64_t size,
bool formatted_buffer_write) {
if (!GuestRange {address, size}.Valid()) {
@@ -1827,37 +1850,6 @@ bool TextureCache::TouchMeta(uint64_t address, uint32_t slice, bool is_clear) {
return true;
}
bool TextureCache::InvalidateMemory(PageFaultAccess access, uint64_t address, uint64_t size,
PageFaultPhase phase) noexcept {
if ((access != PageFaultAccess::Read && access != PageFaultAccess::Write) ||
!GuestRange {address, size}.Valid()) {
return false;
}
if (access == PageFaultAccess::Read) {
return false;
}
if (phase == PageFaultPhase::Invalidate) {
CacheLock lock(*this, m_lock);
const bool tracked =
std::ranges::any_of(FindImagesInRegion(address, size, true), [&](ImageId id) {
const auto owner = ResolveOwner(id);
return owner != nullptr && !owner->depth_id && owner->IsTracked();
});
if (tracked) {
InvalidateCpuAliases(address, size);
}
return tracked;
}
if (phase != PageFaultPhase::Complete && phase != PageFaultPhase::Release) {
return false;
}
CacheLock lock(*this, m_lock);
return std::ranges::any_of(FindImagesInRegion(address, size, true), [&](ImageId id) {
const auto owner = ResolveOwner(id);
return owner != nullptr && !owner->depth_id;
});
}
void TextureCache::UnmapMemory(uint64_t address, uint64_t size) {
if (!GuestRange {address, size}.Valid()) {
EXIT("TextureCache: invalid unmap range\n");
+5 -8
View File
@@ -66,7 +66,6 @@ public:
[[nodiscard]] bool ClearImageFromBuffer(CommandBuffer& command, uint64_t address, uint64_t size,
uint32_t packed_clear);
void InvalidateMemory(uint64_t address, uint64_t size);
[[nodiscard]] bool SynchronizeImageToBuffer(uint64_t address, uint64_t size);
[[nodiscard]] bool InvalidateMemoryFromGPU(uint64_t address, uint64_t size,
bool formatted_buffer_write = false);
[[nodiscard]] RegionInfo QueryRegion(uint64_t address, uint64_t size);
@@ -76,11 +75,9 @@ public:
[[nodiscard]] bool ClearMeta(uint64_t address);
[[nodiscard]] bool TouchMeta(uint64_t address, uint32_t slice, bool is_clear);
[[nodiscard]] bool InvalidateMemory(PageFaultAccess access, uint64_t address, uint64_t size,
PageFaultPhase phase) noexcept;
void UnmapMemory(uint64_t address, uint64_t size);
void ProcessDownloadImages();
void RunGarbageCollector();
void UnmapMemory(uint64_t address, uint64_t size);
void ProcessDownloadImages();
void RunGarbageCollector();
private:
enum class TransferDirection { Upload, Download };
@@ -142,7 +139,7 @@ private:
[[nodiscard]] DownloadPlan BuildDownload(const Image& image) const;
void UploadImage(Image& image, const ImageDesc& desc, Buffer& source, uint64_t source_offset);
void DownloadImageData(Image& image, Buffer& destination, uint64_t destination_offset,
DownloadPlan plan);
uint64_t destination_size, DownloadPlan plan);
void DownloadDepth(Image& image, Buffer& destination, uint64_t destination_offset);
void CommitGpuWrite(Image& image);
void PrepareImageCopy(Image& image);
@@ -157,7 +154,6 @@ private:
void InvalidateCpuAliases(uint64_t address, uint64_t size);
void ClearGpuModified(ImageId id);
[[nodiscard]] bool SynchronizeImageToBuffer(ImageId id);
void DownloadImage(ImageId id);
[[nodiscard]] bool TryDownloadImage(ImageId id);
[[nodiscard]] std::pair<uint8_t*, uint64_t> MapDownload(uint64_t size, uint64_t alignment);
@@ -186,6 +182,7 @@ private:
bool m_readback_linear_images = false;
friend struct TextureCacheTestAccess;
friend class BufferCache;
friend class RenderExecutor;
};
@@ -7,8 +7,8 @@
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/guest_gpu/tile.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
@@ -23,10 +23,10 @@ static std::atomic<uint32_t> g_render_color_log_count = 0;
// NOLINTNEXTLINE(readability-function-cognitive-complexity)
void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandBuffer& buffer,
RenderColorInfo& r,
uint32_t render_target_slice_offset,
uint32_t render_target_slot, bool ignore_target_mask,
bool exact_format) {
RenderColorInfo& r,
uint32_t render_target_slice_offset,
uint32_t render_target_slot, bool ignore_target_mask,
bool exact_format) {
KYTY_PROFILER_FUNCTION();
const auto& hw = buffer.GetRegisters();
@@ -79,10 +79,8 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
const auto view = ResolveTargetViewInfo(
rt.view.base_array_slice_index, rt.view.last_array_slice_index, render_target_slice_offset);
switch (view.type) {
case TargetViewType::Image2D: break;
case TargetViewType::Image2DArray:
EXIT("layered render-target views are unsupported: base=%u count=%u\n", view.base_layer,
view.layer_count);
case TargetViewType::Image2D:
case TargetViewType::Image2DArray: break;
case TargetViewType::Unsupported:
EXIT("invalid render-target view: base=%u last=%u draw_offset=%u\n",
rt.view.base_array_slice_index, rt.view.last_array_slice_index,
@@ -121,7 +119,18 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
uint32_t pitch = 0;
uint64_t size = 0;
bool tile = false;
const bool standard64 =
const bool volume = rt.attrib3.dimension == 2;
if (rt.attrib3.dimension != 1 && !volume) {
EXIT("unsupported render-target dimension: %u\n", rt.attrib3.dimension);
}
if (!volume && rt.attrib3.depth != 0) {
EXIT("2D render target has nonzero depth: %u\n", rt.attrib3.depth);
}
if (volume && samples != 1) {
EXIT("multisampled 3D render targets are unsupported\n");
}
const uint32_t depth = volume ? rt.attrib3.depth + 1u : 1u;
const bool standard64 =
rt.attrib3.tile_mode == Prospero::GpuEnumValue(Prospero::TileMode::kStandard64KB);
switch (rt.attrib3.tile_mode) {
@@ -147,6 +156,7 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
if (bytes_per_element == 0) {
EXIT("render-target format has no valid element size\n");
}
const auto transfer_format = ImageOps::RenderTargetTransferFormat(bytes_per_element);
if (standard64 &&
(rt.attrib3.dimension != 1 || rt.attrib3.depth != 0 || levels != 1 ||
rt.view.current_mip_level != 0 || view.base_layer != 0 || view.image_layers != 1 ||
@@ -167,10 +177,14 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
if (rt.pitch.pitch_div8_minus1 != 0) {
pitch = (rt.pitch.pitch_div8_minus1 + 1u) << 3u;
} else if (tile) {
pitch = standard64
? TileGetTexturePitch(Prospero::GpuEnumValue(Prospero::BufferFormat::k32Float),
width, levels, rt.attrib3.tile_mode)
: TileGetRenderTargetPitch(width, bytes_per_element, rt.attrib.num_fragments);
if (volume) {
pitch = TileGetTexturePitch(transfer_format, width, levels, rt.attrib3.tile_mode);
} else if (standard64) {
pitch = TileGetTexturePitch(Prospero::GpuEnumValue(Prospero::BufferFormat::k32Float),
width, levels, rt.attrib3.tile_mode);
} else {
pitch = TileGetRenderTargetPitch(width, bytes_per_element, rt.attrib.num_fragments);
}
if (pitch == 0) {
EXIT("unsupported render-target pitch: width=%u bytes=%u\n", width, bytes_per_element);
}
@@ -178,9 +192,19 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
pitch = width;
}
TileSizeOffset mip_sizes[16] {};
TilePaddedSize mip_padded[16] {};
if (tile) {
TileSizeOffset mip_sizes[16] {};
TilePaddedSize mip_padded[16] {};
TileVolumeLayout volume_layout {};
uint64_t backing_size = 0;
if (volume) {
if (!tile || !TileGetTextureVolumeLayout(transfer_format, width, height, depth, levels,
rt.attrib3.tile_mode, volume_layout)) {
EXIT("unsupported 3D render-target layout: %ux%ux%u levels=%u tile=%u\n", width, height,
depth, levels, rt.attrib3.tile_mode);
}
size = volume_layout.block_slice_size;
backing_size = volume_layout.total_size;
} else if (tile) {
TileSizeAlign layout {};
bool valid_layout = false;
if (standard64) {
@@ -205,12 +229,6 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
mip_sizes[0] = {static_cast<uint32_t>(size), 0, 0, 0, 0, 0};
mip_padded[0] = {pitch, height};
}
if (rt.slice.slice_div64_minus1 != 0 &&
(static_cast<uint64_t>(rt.slice.slice_div64_minus1) + 1u) * 64u != size) {
EXIT("render-target slice span mismatch: encoded=0x%016" PRIx64 " derived=0x%016" PRIx64
"\n",
(static_cast<uint64_t>(rt.slice.slice_div64_minus1) + 1u) * 64u, size);
}
} else {
size = static_cast<uint64_t>(pitch) * height * bytes_per_element * samples;
if (size > UINT32_MAX) {
@@ -219,23 +237,40 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
mip_sizes[0] = {static_cast<uint32_t>(size), 0, 0, 0, 0, 0};
mip_padded[0] = {pitch, height};
}
if (size == 0 || size > UINT64_MAX / view.image_layers) {
if (rt.slice.slice_div64_minus1 != 0 &&
(static_cast<uint64_t>(rt.slice.slice_div64_minus1) + 1u) * 64u != size) {
EXIT("render-target slice span mismatch: encoded=0x%016" PRIx64 " derived=0x%016" PRIx64
"\n",
(static_cast<uint64_t>(rt.slice.slice_div64_minus1) + 1u) * 64u, size);
}
if (size == 0 || (!volume && size > UINT64_MAX / view.image_layers)) {
EXIT("render-target memory footprint is invalid\n");
}
const auto backing_size = size * view.image_layers;
if (!volume) {
backing_size = size * view.image_layers;
}
if (backing_size == 0) {
EXIT("render-target backing is empty\n");
}
if (backing_size > TRACKER_ADDRESS_SIZE - rt.base.addr) {
EXIT("render-target backing range is invalid\n");
}
const vk::Extent2D view_extent = {std::max(width >> rt.view.current_mip_level, 1u),
std::max(height >> rt.view.current_mip_level, 1u)};
const uint32_t view_depth = std::max(depth >> rt.view.current_mip_level, 1u);
if (volume &&
(view.base_layer >= view_depth || view.layer_count > view_depth - view.base_layer)) {
EXIT("3D render-target view exceeds mip depth: base=%u count=%u depth=%u mip=%u\n",
view.base_layer, view.layer_count, view_depth, rt.view.current_mip_level);
}
auto decision_log_id = g_render_color_log_count.fetch_add(1);
if (decision_log_id < 128) {
LOGF("RenderColorTarget: slot=%" PRIu32 " addr=0x%010" PRIx64 " size=0x%016" PRIx64
" extent=%ux%u view_mip=%u view_extent=%ux%u levels=%u pitch=%u"
" extent=%ux%ux%u view_mip=%u view_extent=%ux%u levels=%u pitch=%u"
" fmt=0x%08" PRIx32 " nfmt=0x%08" PRIx32 " order=0x%08" PRIx32 " samples=%u tile=%s\n",
rt_slot, rt.base.addr, backing_size, width, height, rt.view.current_mip_level,
rt_slot, rt.base.addr, backing_size, width, height, depth, rt.view.current_mip_level,
view_extent.width, view_extent.height, levels, pitch, rt.info.format,
rt.info.channel_type, rt.info.channel_order, samples, tile ? "tiled" : "linear");
}
@@ -244,15 +279,24 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
desc.type = TextureCache::BindingType::RenderTarget;
desc.info.data = {rt.base.addr, backing_size};
desc.info.pixel_format = target_format.format;
desc.info.guest_format = ImageOps::RenderTargetTransferFormat(bytes_per_element);
desc.info.type = Prospero::ImageType::kColor2D;
desc.info.extent = {width, height, 1};
desc.info.resources = {levels, view.image_layers};
desc.info.pitch = pitch;
desc.info.guest_format = transfer_format;
desc.info.type = volume ? Prospero::ImageType::kColor3D : Prospero::ImageType::kColor2D;
desc.info.extent = {width, height, depth};
desc.info.resources = {levels, volume ? 1u : view.image_layers};
desc.info.pitch = pitch;
desc.info.bytes_per_block = bytes_per_element;
desc.info.samples = samples;
desc.info.tile_mode = rt.attrib3.tile_mode;
for (uint32_t level = 0; level < levels; level++) {
if (volume) {
desc.info.mip_layout[level] = {
volume_layout.level_offsets[level],
volume_layout.level_sizes[level],
volume_layout.level_widths[level],
volume_layout.level_heights[level],
};
continue;
}
const auto level_offset =
mip_sizes[level].src_size != 0 ? mip_sizes[level].src_offset : mip_sizes[level].offset;
const auto level_size =
@@ -275,20 +319,20 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
desc.view_info.base_layer = view.base_layer;
desc.view_info.layer_count = view.layer_count;
desc.view_info.usage = vk::ImageUsageFlagBits::eColorAttachment;
auto& texture_cache = m_context.GetTextureCache();
r.desc = std::move(desc);
r.image_id = texture_cache.FindImage(r.desc, exact_format);
r.type = RenderColorType::RenderTexture;
r.base_addr = rt.base.addr;
r.image_view = nullptr;
r.format = r.desc.view_info.format;
r.extent = view_extent;
r.base_mip_level = rt.view.current_mip_level;
r.buffer_size = backing_size;
r.samples = samples;
r.export_mapping = target_format.export_mapping;
r.color_clear_enable = false;
r.color_clear_value = {};
auto& texture_cache = m_context.GetTextureCache();
r.desc = std::move(desc);
r.image_id = texture_cache.FindImage(r.desc, exact_format);
r.type = RenderColorType::RenderTexture;
r.base_addr = rt.base.addr;
r.image_view = nullptr;
r.format = r.desc.view_info.format;
r.extent = view_extent;
r.base_mip_level = rt.view.current_mip_level;
r.buffer_size = backing_size;
r.samples = samples;
r.export_mapping = target_format.export_mapping;
r.color_clear_enable = false;
r.color_clear_value = {};
BindRenderTarget(r.image_id);
}
@@ -2,8 +2,8 @@
#define EMULATOR_SRC_GRAPHICS_HOST_GPU_RENDERER_COLORRENDERTARGET_H_
#include "graphics/guest_gpu/gpu_defs.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include <cstdint>
@@ -44,13 +44,13 @@ CommandSlot* CommandScheduler::CommandPool::CreateSlot() {
allocate.commandPool = m_pool;
allocate.level = vk::CommandBufferLevel::ePrimary;
allocate.commandBufferCount = 1;
vk::CommandBuffer buffer = nullptr;
vk::CommandBuffer buffer = nullptr;
EXIT_IF(graphics.device.allocateCommandBuffers(&allocate, &buffer) != vk::Result::eSuccess);
vk::FenceCreateInfo fence_create {};
fence_create.sType = vk::StructureType::eFenceCreateInfo;
fence_create.flags = vk::FenceCreateFlagBits::eSignaled;
vk::Fence fence = nullptr;
vk::Fence fence = nullptr;
if (graphics.device.createFence(&fence_create, nullptr, &fence) != vk::Result::eSuccess) {
graphics.device.freeCommandBuffers(m_pool, 1, &buffer);
EXIT("failed to create command-buffer fence\n");
@@ -70,9 +70,9 @@ CommandSlot* CommandScheduler::CommandPool::Allocate(GraphicContext& graphics) {
Create(graphics);
}
EXIT_IF(m_graphics != &graphics);
auto found = std::ranges::find_if(m_slots, [](const auto& slot) { return !slot.busy; });
auto* slot = found != m_slots.end() ? &*found : CreateSlot();
slot->busy = true;
auto found = std::ranges::find_if(m_slots, [](const auto& slot) { return !slot.busy; });
auto* slot = found != m_slots.end() ? &*found : CreateSlot();
slot->busy = true;
slot->Reset();
return slot;
}
@@ -331,8 +331,7 @@ void CommandScheduler::WaitPriorityOperations(uint64_t tick) {
EXIT_IF(g_deferred_callback_scheduler == this);
std::unique_lock lock(m_operation_mutex);
m_operation_available.wait(lock, [this, tick] {
const bool active_before_or_at =
m_priority_active && m_priority_active_tick <= tick;
const bool active_before_or_at = m_priority_active && m_priority_active_tick <= tick;
const bool queued_before_or_at =
!m_priority_operations.empty() && m_priority_operations.front().tick <= tick;
return !active_before_or_at && !queued_before_or_at;
@@ -47,21 +47,21 @@ public:
void FinishCurrent();
// Deferred callbacks can observe an externally owned drain, but cannot initiate shutdown:
// the priority runner cannot join itself.
void Shutdown();
void Wait(uint64_t tick);
void PopPendingOperations();
void DrainPriorityOperations();
void WaitPriorityOperations(uint64_t tick);
void DeferOperation(Common::UniqueFunction<void>&& operation);
void DeferPriorityOperation(Common::UniqueFunction<void>&& operation);
void Shutdown();
void Wait(uint64_t tick);
void PopPendingOperations();
void DrainPriorityOperations();
void WaitPriorityOperations(uint64_t tick);
void DeferOperation(Common::UniqueFunction<void>&& operation);
void DeferPriorityOperation(Common::UniqueFunction<void>&& operation);
[[nodiscard]] static bool InDeferredOperation() noexcept;
[[nodiscard]] bool Active() const noexcept { return m_current >= 0; }
void CheckActive() const;
RenderCommandBuffer& Current() const;
[[nodiscard]] uint64_t CurrentTick() const noexcept { return m_master.CurrentTick(); }
[[nodiscard]] bool IsFree(uint64_t tick);
[[nodiscard]] RenderContext& Context() const noexcept { return m_context; }
[[nodiscard]] bool Active() const noexcept { return m_current >= 0; }
void CheckActive() const;
RenderCommandBuffer& Current() const;
[[nodiscard]] uint64_t CurrentTick() const noexcept { return m_master.CurrentTick(); }
[[nodiscard]] bool IsFree(uint64_t tick);
[[nodiscard]] RenderContext& Context() const noexcept { return m_context; }
[[nodiscard]] GraphicContext& Graphics() const noexcept { return m_graphics; }
private:
@@ -91,11 +91,11 @@ private:
uint64_t tick = 0;
};
void BindCurrent() const;
CommandBuffer& SubmitCurrent(SubmitInfo& submit);
void BeginNext();
void PriorityOperationsThread(std::stop_token stop);
void RunOperation(Common::UniqueFunction<void>&& operation);
void BindCurrent() const;
CommandBuffer& SubmitCurrent(SubmitInfo& submit);
void BeginNext();
void PriorityOperationsThread(std::stop_token stop);
void RunOperation(Common::UniqueFunction<void>&& operation);
[[nodiscard]] CommandSlot* AllocateCommandBuffer();
[[nodiscard]] uint64_t NextSubmitSequence() noexcept;
@@ -109,14 +109,14 @@ private:
std::mutex m_operation_mutex;
std::condition_variable m_operation_available;
std::jthread m_priority_thread;
bool m_priority_active = false;
bool m_priority_active = false;
uint64_t m_priority_active_tick = 0;
OperationState m_operation_state = OperationState::Open;
int m_current = -1;
bool m_recording = false;
HW::Context* m_registers = nullptr;
HW::UserConfig* m_user_config = nullptr;
HW::Shader* m_shaders = nullptr;
OperationState m_operation_state = OperationState::Open;
int m_current = -1;
bool m_recording = false;
HW::Context* m_registers = nullptr;
HW::UserConfig* m_user_config = nullptr;
HW::Shader* m_shaders = nullptr;
std::atomic<uint64_t> m_submit_sequence = 0;
friend class CommandBuffer;
+13 -13
View File
@@ -8,8 +8,8 @@
#include "graphics/host_gpu/renderer/colorRenderTarget.h"
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/depthRenderTarget.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/image/imageView.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/vma.h"
@@ -270,30 +270,30 @@ void CommandBuffer::BeginRendering(const RenderState& state) const {
colors[i].sType = vk::StructureType::eRenderingAttachmentInfo;
colors[i].imageView = attachment.image_view;
colors[i].imageLayout = attachment.image_layout;
colors[i].loadOp = attachment.is_clear ? vk::AttachmentLoadOp::eClear
: vk::AttachmentLoadOp::eLoad;
colors[i].storeOp = vk::AttachmentStoreOp::eStore;
colors[i].clearValue.color.uint32 = attachment.clear_value;
colors[i].loadOp =
attachment.is_clear ? vk::AttachmentLoadOp::eClear : vk::AttachmentLoadOp::eLoad;
colors[i].storeOp = vk::AttachmentStoreOp::eStore;
colors[i].clearValue.color.uint32 = attachment.clear_value;
}
const auto& depth_stencil = state.depth_stencil_attachment;
const auto& depth_stencil = state.depth_stencil_attachment;
vk::RenderingAttachmentInfo depth {};
depth.sType = vk::StructureType::eRenderingAttachmentInfo;
depth.imageView = depth_stencil.image_view;
depth.imageLayout = depth_stencil.image_layout;
depth.loadOp = depth_stencil.depth_clear ? vk::AttachmentLoadOp::eClear
: vk::AttachmentLoadOp::eLoad;
depth.storeOp = vk::AttachmentStoreOp::eStore;
depth.loadOp =
depth_stencil.depth_clear ? vk::AttachmentLoadOp::eClear : vk::AttachmentLoadOp::eLoad;
depth.storeOp = vk::AttachmentStoreOp::eStore;
depth.clearValue.depthStencil.depth = std::bit_cast<float>(depth_stencil.clear_value[0]);
vk::RenderingAttachmentInfo stencil {};
stencil.sType = vk::StructureType::eRenderingAttachmentInfo;
stencil.imageView = depth_stencil.image_view;
stencil.imageLayout = depth_stencil.image_layout;
stencil.loadOp = depth_stencil.stencil_clear ? vk::AttachmentLoadOp::eClear
: vk::AttachmentLoadOp::eLoad;
stencil.storeOp = vk::AttachmentStoreOp::eStore;
stencil.clearValue.depthStencil.stencil = depth_stencil.clear_value[1];
stencil.loadOp =
depth_stencil.stencil_clear ? vk::AttachmentLoadOp::eClear : vk::AttachmentLoadOp::eLoad;
stencil.storeOp = vk::AttachmentStoreOp::eStore;
stencil.clearValue.depthStencil.stencil = depth_stencil.clear_value[1];
vk::RenderingInfo rendering {};
rendering.sType = vk::StructureType::eRenderingInfo;
+10 -28
View File
@@ -359,30 +359,12 @@ static void RtCheck(const HW::RenderTarget& rt) {
logged = true;
}
}
if (rt.attrib3.depth != 0x00000000) {
static bool logged = false;
if (!logged) {
LOGF("RenderTarget: temporary: ignoring PS5 color target depth_minus1=0x%08" PRIx32
"\n",
rt.attrib3.depth);
logged = true;
}
}
if (!RenderIsColorTileMode(rt.attrib3.tile_mode)) {
EXIT("unknown PS5 render-target tile mode: 0x%08" PRIx32 "\n", rt.attrib3.tile_mode);
}
if (!RenderIsColorDimension(rt.attrib3.dimension)) {
EXIT("unknown PS5 render-target dimension: 0x%08" PRIx32 "\n", rt.attrib3.dimension);
}
if (rt.attrib3.dimension != 0x00000001) {
static bool logged = false;
if (!logged) {
LOGF("RenderTarget: temporary: using 2D fallback for PS5 color "
"dimension=0x%08" PRIx32 "\n",
rt.attrib3.dimension);
logged = true;
}
}
if (!rt.attrib3.cmask_pipe_aligned) {
static bool logged = false;
if (!logged) {
@@ -497,7 +479,15 @@ static void ZPrint(const char* func, const HW::DepthRenderTarget& z) {
}
// NOLINTNEXTLINE(readability-function-cognitive-complexity)
static void ZCheck(const HW::DepthRenderTarget& z) {
static void ZCheck(const HW::DepthRenderTarget& z, const HW::DepthControl& dc,
const HW::RenderControl& rc) {
const bool depth_active =
dc.z_enable || dc.z_write_enable || dc.depth_bounds_enable || rc.depth_clear_enable;
const bool stencil_active = dc.stencil_enable || rc.stencil_clear_enable;
if (!depth_active && !stencil_active) {
return;
}
EXIT_NOT_IMPLEMENTED(!z.z_info.HasValidTextureCompatibility());
EXIT_NOT_IMPLEMENTED(!z.stencil_info.HasValidTextureCompatibility());
if (z.z_info.format == 0) {
@@ -548,14 +538,6 @@ static void ZCheck(const HW::DepthRenderTarget& z) {
EXIT_NOT_IMPLEMENTED(z.htile_surface.prefetch_height != 0x00000000);
EXIT_NOT_IMPLEMENTED(z.htile_surface.dst_outside_zero_to_one != 0x00000000);
if (z.depth_view.slice_start != 0x00000000 || z.depth_view.slice_max != 0x00000000) {
static std::atomic<uint32_t> log_count {0};
if (log_count.fetch_add(1, std::memory_order_relaxed) < 16) {
LOGF("DepthTarget: temporary: ignoring PS5 array slice view start=0x%08" PRIx32
", max=0x%08" PRIx32 "\n",
z.depth_view.slice_start, z.depth_view.slice_max);
}
}
if (z.depth_view.current_mip_level != 0x00000000) {
static std::atomic<uint32_t> log_count {0};
if (log_count.fetch_add(1, std::memory_order_relaxed) < 16) {
@@ -1214,7 +1196,7 @@ void hw_check(const RenderCommandBuffer& buffer) {
log_phase("vp");
VpCheck(vp, smc);
log_phase("z");
ZCheck(z);
ZCheck(z, d, rc);
log_phase("clip");
ClipCheck(c);
log_phase("rc");
@@ -10,10 +10,10 @@
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/guest_gpu/tile.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/image/imageView.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/vulkanCommon.h"
@@ -150,10 +150,8 @@ void RenderExecutor::ResolveRenderDepthTarget(uint64_t submit_id, RenderCommandB
has_stencil, has_htile, z.stencil_info.htile_stencil_disabled);
const auto view = ResolveTargetViewInfo(z.depth_view.slice_start, z.depth_view.slice_max);
switch (view.type) {
case TargetViewType::Image2D: break;
case TargetViewType::Image2DArray:
DepthFatal("layered depth views are unsupported: base=%u count=%u", view.base_layer,
view.layer_count);
case TargetViewType::Image2D:
case TargetViewType::Image2DArray: break;
case TargetViewType::Unsupported:
DepthFatal("invalid depth view: base=%u last=%u", z.depth_view.slice_start,
z.depth_view.slice_max);
@@ -2,9 +2,9 @@
#define EMULATOR_SRC_GRAPHICS_HOST_GPU_RENDERER_DEPTHRENDERTARGET_H_
#include "common/assert.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/renderer/image/imageView.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include <cstdint>
@@ -182,8 +182,8 @@ void BlitHelper::ReinterpretColorAsMsDepth(Image& source, Image& destination) {
auto command = command_buffer.Handle();
source.Transit(vk::ImageLayout::eShaderReadOnlyOptimal, vk::AccessFlagBits2::eShaderRead, {},
command);
destination.Transit(ColorToMsDepthLayout,
vk::AccessFlagBits2::eDepthStencilAttachmentWrite, {}, command);
destination.Transit(ColorToMsDepthLayout, vk::AccessFlagBits2::eDepthStencilAttachmentWrite, {},
command);
vk::RenderingAttachmentInfo depth_attachment {};
depth_attachment.sType = vk::StructureType::eRenderingAttachmentInfo;
@@ -105,6 +105,10 @@ Image::Barriers Image::GetBarriers(vk::ImageLayout destinat
std::optional<ImageSubresourceRange> range) {
auto& state = backing.state;
auto& subresource_states = backing.subresource_states;
if (range && info.IsVolume()) {
range->base_layer = 0;
range->layer_count = 1;
}
const bool partial =
range && (range->base_level != 0 || range->level_count != info.resources.levels ||
@@ -19,10 +19,9 @@ struct GuestRange {
uint64_t address = 0;
uint64_t size = 0;
[[nodiscard]] constexpr bool Empty() const noexcept { return address == 0 || size == 0; }
[[nodiscard]] constexpr bool Valid() const noexcept {
return !Empty() && address < TRACKER_ADDRESS_SIZE &&
size <= TRACKER_ADDRESS_SIZE - address;
[[nodiscard]] constexpr bool Empty() const noexcept { return address == 0 || size == 0; }
[[nodiscard]] constexpr bool Valid() const noexcept {
return !Empty() && address < TRACKER_ADDRESS_SIZE && size <= TRACKER_ADDRESS_SIZE - address;
}
[[nodiscard]] constexpr uint64_t End() const noexcept { return address + size; }
auto operator<=>(const GuestRange&) const = default;
@@ -47,10 +46,10 @@ struct ImageSubresources {
};
struct ImageSubresourceRange {
uint32_t base_level = 0;
uint32_t level_count = 1;
uint32_t base_layer = 0;
uint32_t layer_count = 1;
uint32_t base_level = 0;
uint32_t level_count = 1;
uint32_t base_layer = 0;
uint32_t layer_count = 1;
auto operator<=>(const ImageSubresourceRange&) const = default;
};
@@ -67,10 +66,10 @@ struct ImageInfo {
GuestRange stencil;
ImageMetadataInfo metadata;
uint32_t htile_clear_mask = UINT32_MAX;
vk::Format pixel_format = vk::Format::eUndefined;
uint32_t guest_format = 0;
Prospero::ImageType type = Prospero::ImageType::kColor2D;
vk::Extent3D extent = {1, 1, 1};
vk::Format pixel_format = vk::Format::eUndefined;
uint32_t guest_format = 0;
Prospero::ImageType type = Prospero::ImageType::kColor2D;
vk::Extent3D extent = {1, 1, 1};
ImageSubresources resources;
uint32_t pitch = 0;
uint32_t bytes_per_block = 0;
@@ -352,8 +351,7 @@ inline bool ImageInfo::IsDepth() const noexcept {
}
const auto transfer_bytes = DepthAspectTransferBytes(info.pixel_format);
return transfer_bytes == info.bytes_per_block ||
(info.bytes_per_block == sizeof(uint16_t) &&
transfer_bytes == sizeof(uint32_t));
(info.bytes_per_block == sizeof(uint16_t) && transfer_bytes == sizeof(uint32_t));
}
[[nodiscard]] inline VideoOutCompression
@@ -470,18 +468,13 @@ IsSupportedDisplayRenderTargetTileMode(uint32_t tile_mode) noexcept {
vk::ClearColorValue& clear) {
vk::ClearColorValue next {};
const auto unorm8 = [](uint32_t value) { return static_cast<float>(value & 0xffu) / 255.0f; };
const auto srgb8 = [](uint32_t value) {
const auto srgb8 = [](uint32_t value) {
const auto encoded = static_cast<float>(value & 0xffu) / 255.0f;
return encoded <= 0.04045f ? encoded / 12.92f
: std::pow((encoded + 0.055f) / 1.055f, 2.4f);
return encoded <= 0.04045f ? encoded / 12.92f : std::pow((encoded + 0.055f) / 1.055f, 2.4f);
};
switch (format) {
case vk::Format::eR32Uint:
next.uint32[0] = packed;
break;
case vk::Format::eR32Sint:
next.int32[0] = static_cast<int32_t>(packed);
break;
case vk::Format::eR32Uint: next.uint32[0] = packed; break;
case vk::Format::eR32Sint: next.int32[0] = static_cast<int32_t>(packed); break;
case vk::Format::eR8G8B8A8Srgb:
next.float32[0] = srgb8(packed);
next.float32[1] = srgb8(packed >> 8u);
@@ -70,15 +70,14 @@ namespace {
}
case vk::ImageType::e3D:
switch (info.type) {
case vk::ImageViewType::e3D:
return info.base_layer == 0 && info.layer_count == 1;
case vk::ImageViewType::e3D: return info.base_layer == 0 && info.layer_count == 1;
case vk::ImageViewType::e2D:
return static_cast<bool>(
image.flags & vk::ImageCreateFlagBits::e2DArrayCompatible) &&
return static_cast<bool>(image.flags &
vk::ImageCreateFlagBits::e2DArrayCompatible) &&
info.level_count == 1 && info.layer_count == 1;
case vk::ImageViewType::e2DArray:
return static_cast<bool>(
image.flags & vk::ImageCreateFlagBits::e2DArrayCompatible) &&
return static_cast<bool>(image.flags &
vk::ImageCreateFlagBits::e2DArrayCompatible) &&
info.level_count == 1;
default: return false;
}
@@ -325,11 +324,10 @@ bool FormatsCompatible(vk::Format base, vk::Format view) noexcept {
} // namespace ImageViewOps
vk::ImageView Image::FindView(const ImageViewInfo& view_info) {
const auto& image = backing;
const auto& image = backing;
auto normalized = view_info;
const bool is_storage =
static_cast<bool>(normalized.usage & vk::ImageUsageFlagBits::eStorage);
normalized.aspect = FullAspectMask(image.format);
const bool is_storage = static_cast<bool>(normalized.usage & vk::ImageUsageFlagBits::eStorage);
normalized.aspect = FullAspectMask(image.format);
if (normalized.aspect & vk::ImageAspectFlagBits::eDepth &&
IsDepthViewFormat(normalized.format)) {
normalized.format = image.format;
@@ -340,28 +338,26 @@ vk::ImageView Image::FindView(const ImageViewInfo& view_info) {
normalized.format = image.format;
normalized.aspect = vk::ImageAspectFlagBits::eStencil;
}
normalized.usage =
is_storage ? vk::ImageUsageFlagBits::eStorage : vk::ImageUsageFlags {};
normalized.usage = is_storage ? vk::ImageUsageFlagBits::eStorage : vk::ImageUsageFlags {};
const bool format_compatible = normalized.format != vk::Format::eUndefined &&
IsCompatibleViewFormat(image.format, normalized.format);
const bool slice_view = image.image_type == vk::ImageType::e3D &&
(normalized.type == vk::ImageViewType::e2D ||
normalized.type == vk::ImageViewType::e2DArray);
const bool slice_view =
image.image_type == vk::ImageType::e3D && (normalized.type == vk::ImageViewType::e2D ||
normalized.type == vk::ImageViewType::e2DArray);
const bool levels_valid = normalized.level_count != 0 &&
normalized.base_level < image.mip_levels &&
normalized.level_count <= image.mip_levels - normalized.base_level;
const auto view_layers = slice_view && levels_valid
? std::max(image.extent.depth >> normalized.base_level, 1u)
: image.layers;
const bool ranges_valid = levels_valid &&
normalized.layer_count != 0 && normalized.base_layer < view_layers &&
const auto view_layers = slice_view && levels_valid
? std::max(image.extent.depth >> normalized.base_level, 1u)
: image.layers;
const bool ranges_valid = levels_valid && normalized.layer_count != 0 &&
normalized.base_layer < view_layers &&
normalized.layer_count <= view_layers - normalized.base_layer;
const bool mapping_valid =
IsComponentSwizzle(normalized.mapping.r) && IsComponentSwizzle(normalized.mapping.g) &&
IsComponentSwizzle(normalized.mapping.b) && IsComponentSwizzle(normalized.mapping.a);
if (image.image == nullptr || !format_compatible || !ranges_valid || !mapping_valid ||
!IsValidViewType(image, normalized) ||
!IsValidAspect(image, normalized.aspect)) {
!IsValidViewType(image, normalized) || !IsValidAspect(image, normalized.aspect)) {
EXIT("invalid image view: image_format=%d view_format=%d type=%d aspect=0x%x "
"mip=%u+%u layer=%u+%u usage=0x%x image_levels=%u image_layers=%u\n",
static_cast<int>(image.format), static_cast<int>(normalized.format),
@@ -88,7 +88,9 @@ SelectSampledDepthView(vk::Format image_format, vk::Format view_format, uint32_t
IsSupportedSampledDepthResource(const ShaderRecompiler::IR::ImageResource& resource) noexcept {
return resource.kind == ShaderRecompiler::IR::ResourceKind::Image &&
(resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2D ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DArray) &&
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DArray ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray) &&
resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::None && resource.read &&
!resource.written && !resource.atomic;
}
@@ -111,6 +113,13 @@ inline void ValidateStorageColorView(vk::Format image_format, vk::Format view_fo
[[nodiscard]] inline bool
IsSupportedStorageImageResource(const ShaderRecompiler::IR::ImageResource& resource) noexcept {
const bool supported_mip =
(resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::None &&
resource.mip_levels == 1u) ||
(resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::DynamicStorage &&
resource.mip_levels > 0u &&
resource.mip_levels <= ShaderRecompiler::IR::ImageResource::MaxMipLevels &&
!resource.read && !resource.atomic);
return (resource.kind == ShaderRecompiler::IR::ResourceKind::StorageImage ||
resource.kind == ShaderRecompiler::IR::ResourceKind::StorageImageUint) &&
(resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim1D ||
@@ -118,7 +127,7 @@ IsSupportedStorageImageResource(const ShaderRecompiler::IR::ImageResource& resou
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2D ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim3D ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DArray) &&
resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::None && resource.written &&
supported_mip && resource.written &&
(!resource.atomic ||
(resource.kind == ShaderRecompiler::IR::ResourceKind::StorageImageUint &&
resource.read)) &&
@@ -70,6 +70,10 @@ constexpr RenderTargetFormatMapping kRenderTargetFormats[] = {
Prospero::ChannelType::kFloat,
Prospero::ChannelOrder::kStandard,
{vk::Format::eB10G11R11UfloatPack32, 4}},
{Prospero::ChannelLayout::k5_6_5,
Prospero::ChannelType::kUNorm,
Prospero::ChannelOrder::kStandard,
{vk::Format::eB5G6R5UnormPack16, 2}},
{Prospero::ChannelLayout::k16,
Prospero::ChannelType::kUNorm,
Prospero::ChannelOrder::kStandard,
@@ -397,10 +401,10 @@ TextureUploadLayout TextureCalcUploadLayout(uint32_t fmt, uint64_t width, uint64
return layout;
}
std::vector<vk::BufferImageCopy>
TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels, bool array_texture,
bool volume_texture) {
std::vector<vk::BufferImageCopy> TextureBuildImageCopies(const TextureUploadLayout& layout,
uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels,
bool array_texture, bool volume_texture) {
uint32_t mip_width = width;
uint32_t mip_height = height;
uint32_t mip_pitch = volume_texture && static_cast<Prospero::TileMode>(layout.tile) !=
@@ -416,14 +420,13 @@ TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint3
const auto mip_depth = GetTextureLevelDepth(depth, i, volume_texture);
for (uint32_t z = 0; z < mip_depth; z++) {
const auto slice_offset = z * layout.slice_stride;
const auto slice_offset = z * layout.slice_stride;
vk::BufferImageCopy region {};
region.bufferOffset =
layout.level_sizes[i].offset + slice_offset;
region.imageSubresource = {vk::ImageAspectFlagBits::eColor, i,
array_texture ? z : 0, 1};
region.imageOffset.z = volume_texture ? static_cast<int>(z) : 0;
region.imageExtent = {mip_width, mip_height, 1};
region.bufferOffset = layout.level_sizes[i].offset + slice_offset;
region.imageSubresource = {vk::ImageAspectFlagBits::eColor, i, array_texture ? z : 0,
1};
region.imageOffset.z = volume_texture ? static_cast<int>(z) : 0;
region.imageExtent = {mip_width, mip_height, 1};
const bool linear =
static_cast<Prospero::TileMode>(layout.tile) == Prospero::TileMode::kLinear;
if (linear) {
@@ -433,9 +436,8 @@ TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint3
const auto align = [](uint32_t value, uint32_t block) {
return ((value + block - 1u) / block) * block;
};
const auto pitch = align(mip_pitch, layout.texel_block);
region.bufferRowLength =
pitch > align(mip_width, layout.texel_block) ? pitch : 0;
const auto pitch = align(mip_pitch, layout.texel_block);
region.bufferRowLength = pitch > align(mip_width, layout.texel_block) ? pitch : 0;
}
regions.push_back(region);
}
@@ -480,8 +482,7 @@ static bool SetGpuTileSize(uint64_t offset, uint64_t length, uint64_t capacity,
return true;
}
bool TextureBuildGpuTileInfos(uint64_t size,
const std::vector<vk::BufferImageCopy>& regions,
bool TextureBuildGpuTileInfos(uint64_t size, const std::vector<vk::BufferImageCopy>& regions,
const TextureUploadLayout& layout, uint32_t fmt, uint32_t depth,
uint64_t levels, std::vector<GpuTileInfo>& out_infos) {
if (size == 0 || levels == 0 || levels > 16 || depth == 0 ||
@@ -522,13 +523,12 @@ bool TextureBuildGpuTileInfos(uint64_t size,
for (uint32_t z = 0; z < mip_depth; z += block.block_depth) {
const uint32_t copy_depth = std::min(block.block_depth, mip_depth - z);
const auto& region = regions[region_base + z];
const auto pitch = region.bufferRowLength != 0
? region.bufferRowLength
: region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
GpuTileInfo info {};
const auto pitch =
region.bufferRowLength != 0 ? region.bufferRowLength : region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
GpuTileInfo info {};
info.family = block.family;
info.bytes_per_element = block.bytes_per_element;
info.linear_offset = region.bufferOffset;
@@ -544,20 +544,17 @@ bool TextureBuildGpuTileInfos(uint64_t size,
return false;
}
info.linear_slice_stride = linear_stride;
info.width = std::max(
(region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max(
(logical_height + element.tall - 1u) / element.tall, 1u);
info.depth = copy_depth;
info.surface_z = block.block_depth == 1
? static_cast<uint32_t>(region.imageOffset.z)
: 0;
info.pitch =
std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.tail_x = tail ? volume.tail_x[level] : 0;
info.tail_y = tail ? volume.tail_y[level] : 0;
info.tail = tail;
info.tiled_width = volume.level_widths[level];
info.width =
std::max((region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max((logical_height + element.tall - 1u) / element.tall, 1u);
info.depth = copy_depth;
info.surface_z =
block.block_depth == 1 ? static_cast<uint32_t>(region.imageOffset.z) : 0;
info.pitch = std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.tail_x = tail ? volume.tail_x[level] : 0;
info.tail_y = tail ? volume.tail_y[level] : 0;
info.tail = tail;
info.tiled_width = volume.level_widths[level];
info.tiled_height = volume.level_heights[level];
infos.push_back(info);
}
@@ -581,12 +578,11 @@ bool TextureBuildGpuTileInfos(uint64_t size,
const auto level_depth = GetTextureLevelDepth(depth, level, layout.volume_texture);
for (uint32_t z = 0; z < level_depth; z++) {
const auto& region = regions[region_index++];
const auto pitch = region.bufferRowLength != 0
? region.bufferRowLength
: region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
const auto pitch =
region.bufferRowLength != 0 ? region.bufferRowLength : region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
GpuTileInfo info {};
info.family = block.family;
info.bytes_per_element = block.bytes_per_element;
@@ -597,16 +593,14 @@ bool TextureBuildGpuTileInfos(uint64_t size,
info.tiled_size)) {
return false;
}
info.width = std::max(
(region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max(
(logical_height + element.tall - 1u) / element.tall, 1u);
info.width =
std::max((region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max((logical_height + element.tall - 1u) / element.tall, 1u);
info.surface_z = base_family == TileBlockFamily::RenderTarget64KB ||
base_family == TileBlockFamily::Depth64KB
? region.imageSubresource.baseArrayLayer
: 0;
info.pitch =
std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.pitch = std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.tail = tail;
info.tail_x = tail ? level_size.x : 0;
info.tail_y = tail ? level_size.y : 0;
@@ -32,20 +32,19 @@ struct TextureUploadLayout {
TilePaddedSize padded_sizes[16] = {};
};
vk::ComponentMapping TextureGetComponentMapping(uint32_t swizzle);
vk::ComponentMapping TextureGetComponentMapping(uint32_t swizzle);
vk::Format TextureGetFormat(uint32_t fmt);
RenderTargetFormatInfo TextureGetRenderTargetFormat(uint32_t layout, uint32_t type, uint32_t order);
TextureUploadLayout TextureCalcUploadLayout(uint32_t fmt, uint64_t width, uint64_t height,
uint64_t levels, uint32_t depth, uint64_t pitch,
uint64_t tile, uint64_t upload_size,
bool allow_depth_tile, bool volume_texture,
const char* owner);
std::vector<vk::BufferImageCopy>
TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels, bool array_texture,
bool volume_texture);
bool TextureBuildGpuTileInfos(uint64_t size,
const std::vector<vk::BufferImageCopy>& regions,
TextureUploadLayout TextureCalcUploadLayout(uint32_t fmt, uint64_t width, uint64_t height,
uint64_t levels, uint32_t depth, uint64_t pitch,
uint64_t tile, uint64_t upload_size,
bool allow_depth_tile, bool volume_texture,
const char* owner);
std::vector<vk::BufferImageCopy> TextureBuildImageCopies(const TextureUploadLayout& layout,
uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels,
bool array_texture, bool volume_texture);
bool TextureBuildGpuTileInfos(uint64_t size, const std::vector<vk::BufferImageCopy>& regions,
const TextureUploadLayout& layout, uint32_t fmt, uint32_t depth,
uint64_t levels, std::vector<GpuTileInfo>& infos);
@@ -14,9 +14,9 @@
#include "gpu_tiler_shaders/gpu_tiler_standard64_spv.h"
#include "gpu_tiler_shaders/gpu_tiler_swap_bgra16_spv.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/cache/streamBuffer.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
#include "graphics/host_gpu/renderer/image/image.h"
#include "graphics/host_gpu/renderer/cache/streamBuffer.h"
#include <algorithm>
#include <array>
@@ -26,8 +26,8 @@ MasterSemaphore::~MasterSemaphore() {
}
void MasterSemaphore::Refresh() {
uint64_t counter = 0;
const auto result = m_graphics.device.getSemaphoreCounterValue(m_semaphore, &counter);
uint64_t counter = 0;
const auto result = m_graphics.device.getSemaphoreCounterValue(m_semaphore, &counter);
EXIT_NOT_IMPLEMENTED(result != vk::Result::eSuccess);
auto known = m_gpu_tick.load(std::memory_order_acquire);
@@ -22,7 +22,7 @@ public:
[[nodiscard]] uint64_t KnownGpuTick() const noexcept {
return m_gpu_tick.load(std::memory_order_acquire);
}
[[nodiscard]] bool IsFree(uint64_t tick) const noexcept { return KnownGpuTick() >= tick; }
[[nodiscard]] bool IsFree(uint64_t tick) const noexcept { return KnownGpuTick() >= tick; }
[[nodiscard]] uint64_t NextTick() noexcept {
return m_current_tick.fetch_add(1, std::memory_order_release);
}
@@ -26,11 +26,15 @@ bool IsSampledImage(BindingKind kind) {
case BindingKind::Sampled1DArray:
case BindingKind::Sampled2D:
case BindingKind::Sampled2DArray:
case BindingKind::Sampled2DMsaa:
case BindingKind::Sampled2DMsaaArray:
case BindingKind::Sampled3D:
case BindingKind::SampledUint1D:
case BindingKind::SampledUint1DArray:
case BindingKind::SampledUint2D:
case BindingKind::SampledUint2DArray:
case BindingKind::SampledUint2DMsaa:
case BindingKind::SampledUint2DMsaaArray:
case BindingKind::SampledUint3D: return true;
default: return false;
}
@@ -90,10 +94,11 @@ vk::DescriptorBufferInfo BufferInfo(const BufferView& view) {
} // namespace
vk::DescriptorImageInfo DescriptorCache::MakeImageInfo(const TextureBinding& texture) {
EXIT_IF(!texture.image_id || texture.image_view == nullptr ||
texture.layout == vk::ImageLayout::eUndefined);
return {nullptr, texture.image_view, texture.layout};
vk::DescriptorImageInfo DescriptorCache::MakeImageInfo(const TextureBinding& texture,
uint32_t mip) {
const auto view = texture.mip_views.empty() ? texture.image_view : texture.mip_views.at(mip);
EXIT_IF(!texture.image_id || view == nullptr || texture.layout == vk::ImageLayout::eUndefined);
return {nullptr, view, texture.layout};
}
DescriptorCache::~DescriptorCache() {
@@ -150,7 +155,8 @@ void DescriptorCache::CreatePool() {
MaxSets * (ShaderRecompiler::IR::ShaderInfo::MaxBuffers +
ShaderRecompiler::IR::ShaderInfo::MaxAddresses + 3u)},
{vk::DescriptorType::eSampledImage, MaxSets * ShaderRecompiler::IR::ShaderInfo::MaxImages},
{vk::DescriptorType::eStorageImage, MaxSets * ShaderRecompiler::IR::ShaderInfo::MaxImages},
{vk::DescriptorType::eStorageImage, MaxSets * ShaderRecompiler::IR::ShaderInfo::MaxImages *
ShaderRecompiler::IR::ImageResource::MaxMipLevels},
{vk::DescriptorType::eSampler, MaxSets * ShaderRecompiler::IR::ShaderInfo::MaxSamplers},
};
vk::DescriptorPoolCreateInfo info {};
@@ -225,8 +231,10 @@ VulkanDescriptorSet& DescriptorCache::GetDescriptor(Stage
auto* set = Allocate(stage, program);
EXIT_NOT_IMPLEMENTED(set == nullptr);
const auto descriptor_count = program.info.buffers.size() + program.info.images.size() +
program.info.samplers.size() + program.info.addresses.size() + 3u;
uint32_t descriptor_count = 0;
for (const auto& binding: program.bindings.descriptors) {
descriptor_count += DescriptorCount(binding);
}
std::vector<vk::DescriptorBufferInfo> buffer_infos;
std::vector<vk::DescriptorImageInfo> image_infos;
std::vector<vk::WriteDescriptorSet> writes;
@@ -234,6 +242,7 @@ VulkanDescriptorSet& DescriptorCache::GetDescriptor(Stage
image_infos.reserve(descriptor_count);
writes.reserve(program.bindings.descriptors.size());
std::vector<uint32_t> image_mips(program.info.images.size());
for (const auto& binding: program.bindings.descriptors) {
vk::WriteDescriptorSet write {};
write.sType = vk::StructureType::eWriteDescriptorSet;
@@ -269,7 +278,11 @@ VulkanDescriptorSet& DescriptorCache::GetDescriptor(Stage
default: {
for (const auto resource: binding.resources) {
const auto& texture = data.images.at(resource);
image_infos.push_back(MakeImageInfo(texture));
const auto mip = program.info.images.at(resource).mip_mode ==
ShaderRecompiler::IR::ImageMipMode::DynamicStorage
? image_mips.at(resource)++
: 0u;
image_infos.push_back(MakeImageInfo(texture, mip));
}
break;
}
@@ -47,10 +47,11 @@ public:
enum class Stage { Unknown, Vertex, Pixel, Compute };
struct TextureBinding {
ImageId image_id;
vk::ImageView image_view = nullptr;
TextureCache::ImageDesc desc;
vk::ImageLayout layout = vk::ImageLayout::eUndefined;
ImageId image_id;
vk::ImageView image_view = nullptr;
TextureCache::ImageDesc desc;
vk::ImageLayout layout = vk::ImageLayout::eUndefined;
std::vector<vk::ImageView> mip_views;
};
struct NativeDescriptors {
@@ -94,8 +95,8 @@ private:
int next_free_pool = -1;
};
static vk::DescriptorImageInfo MakeImageInfo(const TextureBinding& texture);
void CreatePool();
static vk::DescriptorImageInfo MakeImageInfo(const TextureBinding& texture, uint32_t mip = 0);
void CreatePool();
VulkanDescriptorSet* Allocate(Stage stage, const ShaderRecompiler::IR::Program& program);
vk::DescriptorSetLayout
GetDescriptorSetLayoutInternal(Stage stage, const ShaderRecompiler::IR::Program& program);
@@ -73,6 +73,11 @@ static Prospero::ImageType TextureBaseType(Prospero::ImageType type) {
}
}
static bool IsMultisampledTexture(Prospero::ImageType type) {
return type == Prospero::ImageType::kColor2DMsaa ||
type == Prospero::ImageType::kColor2DMsaaArray;
}
static BufferView NativeStorageBuffer(RenderContext& context, CommandBuffer& command_buffer,
const ShaderBufferResource& descriptor,
const ShaderRecompiler::IR::BufferResource& resource,
@@ -159,6 +164,8 @@ static bool IsSupportedSampledColorResource(const ShaderRecompiler::IR::ImageRes
case ShaderRecompiler::Decoder::ImageDimension::Dim1DArray:
case ShaderRecompiler::Decoder::ImageDimension::Dim2D:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DArray:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray:
supported_dimension = true;
break;
default: break;
@@ -195,6 +202,22 @@ TargetTextureViewInfo ResolveTargetTextureView(const ShaderRecompiler::IR::Image
? TargetTextureViewInfo {vk::ImageViewType::e2DArray, base_layer,
image_layers - base_layer}
: TargetTextureViewInfo {};
case Prospero::ImageType::kColor2DMsaa:
return resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa &&
base_layer == 0 && image_layers == 1
? TargetTextureViewInfo {vk::ImageViewType::e2D, 0, 1}
: TargetTextureViewInfo {};
case Prospero::ImageType::kColor2DMsaaArray:
if (resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa &&
base_layer == 0 && image_layers == 1) {
return {vk::ImageViewType::e2D, 0, 1};
}
return resource.dimension ==
ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray &&
base_layer < image_layers
? TargetTextureViewInfo {vk::ImageViewType::e2DArray, base_layer,
image_layers - base_layer}
: TargetTextureViewInfo {};
default: return {};
}
}
@@ -209,41 +232,64 @@ bool IsSupportedSampledVideoOutView(const ShaderRecompiler::IR::ImageResource& r
}
bool IsSupportedDepthTargetDescriptor(const ShaderTextureResource& descriptor, const Image& image) {
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto pitch = TileGetTexturePitch(descriptor.Format(), width, 1, descriptor.TileMode());
const auto type = static_cast<Prospero::ImageType>(descriptor.Type());
const bool supported_single_layer =
image.info.resources.layers == 1 && descriptor.Depth() == 0 &&
descriptor.BaseArray5() == 0 &&
(type == Prospero::ImageType::kColor2D || type == Prospero::ImageType::kColor2DArray);
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto type = static_cast<Prospero::ImageType>(descriptor.Type());
const bool multisampled = IsMultisampledTexture(type);
const auto samples = multisampled ? 1u << descriptor.LastLevel() : 1u;
const auto pitch =
multisampled ? TileGetDepthPitch(width, image.info.bytes_per_block, descriptor.LastLevel())
: TileGetTexturePitch(descriptor.Format(), width, 1, descriptor.TileMode());
const bool supported_2d = type == Prospero::ImageType::kColor2D &&
image.info.resources.layers == 1 && descriptor.Depth() == 0 &&
descriptor.BaseArray5() == 0;
const bool supported_array = type == Prospero::ImageType::kColor2DArray &&
descriptor.BaseArray5() <= descriptor.Depth() &&
descriptor.Depth() < image.info.resources.layers;
const bool supported_cube =
type == Prospero::ImageType::kCube && width == height && image.info.resources.layers >= 6 &&
image.info.resources.layers % 6u == 0 &&
static_cast<uint32_t>(descriptor.Depth()) + 1u == image.info.resources.layers &&
descriptor.BaseArray5() == 0;
const bool supported_msaa_2d = type == Prospero::ImageType::kColor2DMsaa &&
image.info.resources.layers == 1 && descriptor.Depth() == 0 &&
descriptor.BaseArray5() == 0;
const bool supported_msaa_array = type == Prospero::ImageType::kColor2DMsaaArray &&
descriptor.BaseArray5() <= descriptor.Depth() &&
descriptor.Depth() < image.info.resources.layers;
const bool levels_ok =
multisampled
? descriptor.BaseLevel() == 0 && descriptor.LastLevel() >= 1 &&
descriptor.LastLevel() <= 3 && descriptor.MaxMip() == descriptor.LastLevel() &&
image.info.resources.levels == 1 && image.info.samples == samples
: descriptor.BaseLevel() == 0 && descriptor.LastLevel() == 0 &&
descriptor.MaxMip() == 0 && image.info.samples == 1;
return image.info.IsDepth() && width == image.info.extent.width &&
height == image.info.extent.height && (supported_single_layer || supported_cube) &&
descriptor.BaseLevel() == 0 && descriptor.LastLevel() == 0 && descriptor.MaxMip() == 0 &&
descriptor.MinLod() == 0 && descriptor.BaseArray5() == 0 &&
height == image.info.extent.height &&
(supported_2d || supported_array || supported_cube || supported_msaa_2d ||
supported_msaa_array) &&
levels_ok && descriptor.MinLod() == 0 &&
descriptor.TileMode() == Prospero::GpuEnumValue(Prospero::TileMode::kDepth) &&
descriptor.BCSwizzle() == 0 && !descriptor.MsaaDepth() && pitch >= width &&
pitch == image.info.pitch;
descriptor.BCSwizzle() == 0 && (!descriptor.MsaaDepth() || multisampled) &&
pitch >= width && pitch == image.info.pitch;
}
bool IsSupportedDepthTextureEncoding(const ShaderTextureResource& descriptor, const Image& image) {
constexpr uint32_t field1_reserved_mask = 0x200fff00u;
constexpr uint32_t field2_reserved_mask = 0xf0003000u;
constexpr uint32_t field3_common = 0x01800000u;
constexpr uint32_t field5_expected = 0x00700000u;
const uint32_t field3_expected =
(descriptor.Type() << 28u) | field3_common | descriptor.DstSelXYZW();
const uint32_t field4_expected = descriptor.Depth() | (descriptor.BaseArray5() << 16u);
const bool common = (descriptor.fields[1] & field1_reserved_mask) == 0 &&
(descriptor.fields[2] & field2_reserved_mask) == 0 &&
descriptor.fields[3] == field3_expected &&
descriptor.fields[4] == field4_expected &&
descriptor.fields[5] == field5_expected;
const uint32_t field3_expected = descriptor.DstSelXYZW() |
(static_cast<uint32_t>(descriptor.BaseLevel()) << 12u) |
(static_cast<uint32_t>(descriptor.LastLevel()) << 16u) |
(static_cast<uint32_t>(descriptor.TileMode()) << 20u) |
(static_cast<uint32_t>(descriptor.Type()) << 28u);
const uint32_t field4_expected = descriptor.Depth() | (descriptor.BaseArray5() << 16u);
const uint32_t field5_expected =
0x00700000u | (static_cast<uint32_t>(descriptor.MaxMip()) << 4u);
const bool common = (descriptor.fields[1] & field1_reserved_mask) == 0 &&
(descriptor.fields[2] & field2_reserved_mask) == 0 &&
descriptor.fields[3] == field3_expected &&
descriptor.fields[4] == field4_expected &&
descriptor.fields[5] == field5_expected;
if (!common || (descriptor.fields[6] == 0 && descriptor.fields[7] != 0)) {
return false;
}
@@ -251,8 +297,9 @@ bool IsSupportedDepthTextureEncoding(const ShaderTextureResource& descriptor, co
return true;
}
constexpr uint32_t htile_control = 0x00280000u;
const auto metadata_addr = descriptor.MetaAddr() << 8u;
return (descriptor.fields[6] & 0x00ffffffu) == htile_control && metadata_addr != 0 &&
const uint32_t expected_control = htile_control | (descriptor.MsaaDepth() ? (1u << 10u) : 0u);
const auto metadata_addr = descriptor.MetaAddr() << 8u;
return (descriptor.fields[6] & 0x00ffffffu) == expected_control && metadata_addr != 0 &&
descriptor.TileMode() == Prospero::GpuEnumValue(Prospero::TileMode::kDepth) &&
image.info.tile_mode == Prospero::GpuEnumValue(Prospero::TileMode::kDepth) &&
image.info.metadata.kind == ImageMetadataKind::Htile &&
@@ -344,9 +391,16 @@ static bool IsSupportedStorageTextureDescriptor(const ShaderRecompiler::IR::Imag
const bool supported_swizzle =
IsValidImageSwizzle(swizzle) &&
(swizzle == DstSel(4, 5, 6, 7) || !resource.read || resource.atomic);
const auto base_level = static_cast<uint32_t>(descriptor.BaseLevel());
const auto last_level = static_cast<uint32_t>(descriptor.LastLevel());
const auto mip_levels = last_level >= base_level ? last_level - base_level + 1u : 0u;
const bool dynamic_mip =
resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::DynamicStorage;
const bool supported_mip_view = descriptor.BaseLevel() == 0 || is_1d || is_2d;
return (is_1d || is_1d_array || is_2d || is_2d_array || is_3d) && supported_tile &&
supported_mip_view && descriptor.BaseLevel() == descriptor.LastLevel() &&
supported_mip_view && mip_levels != 0u &&
((dynamic_mip && mip_levels == resource.mip_levels) ||
(!dynamic_mip && descriptor.BaseLevel() == descriptor.LastLevel())) &&
descriptor.LastLevel() <= descriptor.MaxMip() && descriptor.MinLod() == 0 &&
supported_swizzle && descriptor.BCSwizzle() == 0 && !descriptor.MsaaDepth();
}
@@ -377,10 +431,14 @@ void ValidateStorageTexture(const ShaderRecompiler::IR::ImageResource& resource,
const bool encoding_ok = IsSupportedStorageTextureEncoding(descriptor);
const bool uint_resource =
resource.kind == ShaderRecompiler::IR::ResourceKind::StorageImageUint;
const bool raw_sint_storage =
format == Prospero::GpuEnumValue(Prospero::BufferFormat::k32SInt) && uint_resource &&
resource.written && !resource.read && !resource.atomic;
const bool format_ok =
Prospero::IsSupportedTextureFormat(format) &&
uint_resource == Prospero::IsUintTextureFormat(format) &&
(!resource.atomic || format == Prospero::GpuEnumValue(Prospero::BufferFormat::k32UInt));
raw_sint_storage ||
(Prospero::IsSupportedTextureFormat(format) &&
uint_resource == Prospero::IsUintTextureFormat(format) &&
(!resource.atomic || format == Prospero::GpuEnumValue(Prospero::BufferFormat::k32UInt)));
if (resource_ok && descriptor_ok && encoding_ok && format_ok && size != 0) {
return;
}
@@ -518,6 +576,7 @@ static ImageViewInfo TextureViewInfo(const ShaderRecompiler::IR::ImageResource&
view.layer_count = 1;
break;
case ShaderRecompiler::Decoder::ImageDimension::Dim2DArray:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray:
view.type = vk::ImageViewType::e2DArray;
view.base_layer = descriptor.BaseArray5();
if (view.base_layer >= image_layers) {
@@ -526,6 +585,7 @@ static ImageViewInfo TextureViewInfo(const ShaderRecompiler::IR::ImageResource&
view.layer_count = image_layers - view.base_layer;
break;
case ShaderRecompiler::Decoder::ImageDimension::Dim2D:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa:
view.type = vk::ImageViewType::e2D;
view.base_layer = descriptor.BaseArray5();
if (view.base_layer >= image_layers) {
@@ -556,29 +616,41 @@ RenderExecutor::ResolveTexture(const ShaderRecompiler::IR::ImageResource& reso
return {id, nullptr, std::move(desc)};
}
const auto address = descriptor.Base40();
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto base_level = descriptor.BaseLevel();
const auto last_level = descriptor.LastLevel();
const auto type = TextureType(descriptor);
const bool multisampled =
type == Prospero::ImageType::kColor2DMsaa || type == Prospero::ImageType::kColor2DMsaaArray;
const auto levels = multisampled ? 1u : static_cast<uint32_t>(descriptor.MaxMip()) + 1u;
const auto tile = descriptor.TileMode();
const bool msaa_tile = tile == Prospero::GpuEnumValue(Prospero::TileMode::kRenderTarget);
const auto address = descriptor.Base40();
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto base_level = descriptor.BaseLevel();
const auto last_level = descriptor.LastLevel();
const auto type = TextureType(descriptor);
const bool multisampled = IsMultisampledTexture(type);
const auto levels = multisampled ? 1u : static_cast<uint32_t>(descriptor.MaxMip()) + 1u;
const auto tile = descriptor.TileMode();
const bool depth_tile = tile == Prospero::GpuEnumValue(Prospero::TileMode::kDepth);
const bool msaa_tile =
depth_tile || tile == Prospero::GpuEnumValue(Prospero::TileMode::kRenderTarget);
const bool msaa_array = type == Prospero::ImageType::kColor2DMsaaArray;
if ((!multisampled && (base_level > last_level || last_level >= levels)) ||
(multisampled &&
(base_level != 0 || last_level == 0 || last_level > 3 ||
descriptor.MaxMip() != last_level || !msaa_tile || descriptor.MsaaDepth() ||
descriptor.MaxMip() != last_level || !msaa_tile ||
(descriptor.MsaaDepth() && !depth_tile) ||
(!msaa_array && (descriptor.Depth() != 0 || descriptor.BaseArray5() != 0))))) {
EXIT("unsupported texture mip view: base=%u last=%u levels=%u\n", base_level, last_level,
levels);
EXIT("unsupported texture mip view: base=%u last=%u levels=%u max=%u type=%u tile=%u "
"kind=%u dimension=%u mip_mode=%u read=%d written=%d "
"dwords=%08x,%08x,%08x,%08x,%08x,%08x,%08x,%08x\n",
base_level, last_level, levels, descriptor.MaxMip(), descriptor.Type(), tile,
static_cast<uint32_t>(resource.kind), static_cast<uint32_t>(resource.dimension),
static_cast<uint32_t>(resource.mip_mode), resource.read, resource.written,
descriptor.fields[0], descriptor.fields[1], descriptor.fields[2], descriptor.fields[3],
descriptor.fields[4], descriptor.fields[5], descriptor.fields[6],
descriptor.fields[7]);
}
const auto samples = multisampled ? 1u << last_level : 1u;
const auto view_levels =
multisampled ? 1u : static_cast<uint32_t>(last_level - base_level) + 1u;
multisampled ||
(storage && resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::DynamicStorage)
? 1u
: static_cast<uint32_t>(last_level - base_level) + 1u;
const auto depth = static_cast<uint32_t>(descriptor.Depth()) + 1u;
const auto format = descriptor.Format();
const bool sampled_numeric_class =
@@ -601,7 +673,8 @@ RenderExecutor::ResolveTexture(const ShaderRecompiler::IR::ImageResource& reso
TileSizeAlign size {};
if (multisampled) {
const auto bytes = Prospero::NumBytesPerElement(format);
pitch = TileGetRenderTargetPitch(width, bytes, last_level);
pitch = depth_tile ? TileGetDepthPitch(width, bytes, last_level)
: TileGetRenderTargetPitch(width, bytes, last_level);
if (pitch == 0 || !TileGetRenderTargetSize(width, height, pitch, bytes, size, last_level) ||
size.size > UINT32_MAX / image_layers) {
EXIT("unsupported multisample texture layout\n");
@@ -616,12 +689,13 @@ RenderExecutor::ResolveTexture(const ShaderRecompiler::IR::ImageResource& reso
(address & (static_cast<uint64_t>(size.align) - 1u)) != 0);
if (storage) {
ValidateStorageTexture(resource, descriptor, size.size);
m_context.GetBufferCache().ValidateGpuAccess(address, size.size, resource.read,
resource.written);
}
const auto pixel_format = TextureGetFormat(format);
const auto storage_view_format = SrgbStorageViewFormat(pixel_format);
const auto pixel_format = TextureGetFormat(format);
const auto storage_view_format =
storage && format == Prospero::GpuEnumValue(Prospero::BufferFormat::k32SInt)
? vk::Format::eR32Uint
: SrgbStorageViewFormat(pixel_format);
const auto view_format = storage && storage_view_format != vk::Format::eUndefined
? storage_view_format
: pixel_format;
@@ -831,7 +905,17 @@ void RenderExecutor::RebindImages(CommandBuffer& buffer,
}
auto& binding = images[i];
binding.image_view = texture_cache.FindTexture(binding.image_id, binding.desc);
auto& image = texture_cache.GetImage(binding.image_id);
auto& image = texture_cache.GetImage(binding.image_id);
binding.mip_views.clear();
if (program.info.images[i].mip_mode == ShaderRecompiler::IR::ImageMipMode::DynamicStorage) {
binding.mip_views.reserve(program.info.images[i].mip_levels);
for (uint32_t mip = 0; mip < program.info.images[i].mip_levels; mip++) {
auto view = binding.desc.view_info;
view.base_level += mip;
view.level_count = 1;
binding.mip_views.push_back(mip == 0 ? binding.image_view : image.FindView(view));
}
}
const bool storage = binding.desc.type == TextureCache::BindingType::Storage;
image.usage.storage |= storage;
image.usage.texture |= !storage;
@@ -909,7 +993,11 @@ void RenderExecutor::CommitBindings(CommandBuffer& buffer,
auto& image = m_context.GetTextureCache().GetImage(descriptors.images[i].image_id);
auto& binding = descriptors.images[i];
const auto& view = binding.desc.view_info;
const ImageSubresourceRange range {view.base_level, view.level_count, view.base_layer,
const auto level_count =
program.info.images[i].mip_mode == ShaderRecompiler::IR::ImageMipMode::DynamicStorage
? program.info.images[i].mip_levels
: view.level_count;
const ImageSubresourceRange range {view.base_level, level_count, view.base_layer,
view.layer_count};
const bool storage = binding.desc.type == TextureCache::BindingType::Storage;
if (image.info.data.Empty()) {
@@ -36,7 +36,7 @@ ResolveTargetTextureView(const ShaderRecompiler::IR::ImageResource& resource,
[[nodiscard]] bool IsSupportedDepthTargetDescriptor(const ShaderTextureResource& descriptor,
const Image& image);
[[nodiscard]] bool IsSupportedDepthTextureEncoding(const ShaderTextureResource& descriptor,
const Image& image);
const Image& image);
[[nodiscard]] bool
IsSupportedSampledVideoOutView(const ShaderRecompiler::IR::ImageResource& resource,
const ShaderTextureResource& descriptor, const Image& image);
@@ -88,12 +88,12 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
PipelineStaticParameters static_params {};
GraphicsPipeline p {};
p.ps_shader_id = ps_id;
p.vs_shader_id = vs_id;
p.ps_shader_id = ps_id;
p.vs_shader_id = vs_id;
static_params.color_count = color_count;
PipelineRenderingState rendering {};
rendering.color_count = color_count;
rendering.color_count = color_count;
uint32_t attachment_samples = 0;
for (uint32_t i = 0; i < color_count; i++) {
EXIT_IF(!colors[i].image_id || colors[i].format == vk::Format::eUndefined);
@@ -116,8 +116,8 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
if (attachment_samples == 0) {
attachment_samples = depth.samples;
} else if (attachment_samples != depth.samples) {
EXIT("mixed color/depth sample counts are unsupported: %u and %u\n",
attachment_samples, depth.samples);
EXIT("mixed color/depth sample counts are unsupported: %u and %u\n", attachment_samples,
depth.samples);
}
}
EXIT_IF(attachment_samples == 0 ||
@@ -179,10 +179,10 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
NormalizeStaticParamsForDynamicState(static_params);
GraphicsPipelineKey key {};
key.rendering = rendering;
key.vs_shader_id = p.vs_shader_id;
key.ps_shader_id = p.ps_shader_id;
key.static_params = static_params;
key.rendering = rendering;
key.vs_shader_id = p.vs_shader_id;
key.ps_shader_id = p.ps_shader_id;
key.static_params = static_params;
if (auto iter = m_graphics_pipelines.find(key); iter != m_graphics_pipelines.end()) {
return *iter->second;
@@ -203,9 +203,8 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
LogPipelineTrace("CreatePipelineInternal begin", vs_id.hash0, vs_id.crc32, ps_id.hash0,
ps_id.crc32);
CreatePipelineInternal(m_graphics, m_descriptor_cache, *cached, rendering, vs_input_info,
vs_spirv, ps_input_info,
ps_spirv, static_params, vs_id.hash0, vs_id.crc32, ps_id.hash0,
ps_id.crc32, ps_active);
vs_spirv, ps_input_info, ps_spirv, static_params, vs_id.hash0,
vs_id.crc32, ps_id.hash0, ps_id.crc32, ps_active);
LogPipelineTrace("CreatePipelineInternal done", vs_id.hash0, vs_id.crc32, ps_id.hash0,
ps_id.crc32);
@@ -88,9 +88,9 @@ static_assert(sizeof(PipelineStaticParameters) ==
struct PipelineRenderingState {
std::array<vk::Format, RENDER_COLOR_ATTACHMENTS_MAX> color_formats {};
vk::Format depth_format = vk::Format::eUndefined;
vk::Format stencil_format = vk::Format::eUndefined;
uint32_t color_count = 0;
vk::Format depth_format = vk::Format::eUndefined;
vk::Format stencil_format = vk::Format::eUndefined;
uint32_t color_count = 0;
bool operator==(const PipelineRenderingState&) const = default;
};
@@ -118,11 +118,12 @@ public:
ShaderId cs_shader_id;
};
GraphicsPipeline& CreateGraphicsPipeline(
RenderColorInfo* colors, uint32_t color_count, RenderDepthInfo& depth,
ShaderVertexInputInfo& vs_input_info, RenderCommandBuffer& command,
ShaderPixelInputInfo* ps_input_info, vk::PrimitiveTopology topology, bool ps_active,
std::span<const uint32_t> vs_spirv, std::span<const uint32_t> ps_spirv);
GraphicsPipeline&
CreateGraphicsPipeline(RenderColorInfo* colors, uint32_t color_count, RenderDepthInfo& depth,
ShaderVertexInputInfo& vs_input_info, RenderCommandBuffer& command,
ShaderPixelInputInfo* ps_input_info, vk::PrimitiveTopology topology,
bool ps_active, std::span<const uint32_t> vs_spirv,
std::span<const uint32_t> ps_spirv);
ComputePipeline& CreateComputePipeline(ShaderComputeInputInfo& input_info,
const HW::ComputeShaderInfo& cs_regs,
std::span<const uint32_t> cs_spirv);
@@ -199,7 +200,7 @@ private:
}
};
GraphicContext& m_graphics;
GraphicContext& m_graphics;
DescriptorCache& m_descriptor_cache;
std::unordered_map<GraphicsPipelineKey, std::unique_ptr<GraphicsPipeline>,
GraphicsPipelineKeyHash>
@@ -211,16 +212,13 @@ private:
void LogPipelineTrace(const char* phase, uint32_t vs_hash0, uint32_t vs_crc32, uint32_t ps_hash0,
uint32_t ps_crc32);
void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline,
const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info,
std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info,
std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0,
uint32_t vs_crc32, uint32_t ps_hash0, uint32_t ps_crc32,
bool ps_active);
void CreatePipelineInternal(
GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline, const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info, std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info, std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0, uint32_t vs_crc32,
uint32_t ps_hash0, uint32_t ps_crc32, bool ps_active);
void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::ComputePipeline& pipeline,
const ShaderComputeInputInfo& input_info,
@@ -8,10 +8,10 @@
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "graphics/host_gpu/renderer/pipeline/shaderSubgroup.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/renderer/pipeline/shaderSubgroup.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include "graphics/shader/recompiler/ir/ShaderIR.h"
#include "graphics/shader/shader.h"
@@ -385,9 +385,8 @@ static vk::BlendOp GetBlendOp(uint32_t op) {
return vk::BlendOp::eAdd;
}
static void CreateLayout(DescriptorCache& descriptor_cache,
std::span<vk::DescriptorSetLayout> set_layouts,
uint32_t& set_layouts_num,
static void CreateLayout(DescriptorCache& descriptor_cache,
std::span<vk::DescriptorSetLayout> set_layouts, uint32_t& set_layouts_num,
std::span<vk::PushConstantRange> push_constant_info,
uint32_t& push_constant_info_num,
const ShaderRecompiler::IR::Program& program,
@@ -412,12 +411,11 @@ static void CreateLayout(DescriptorCache& descriptor_cache,
}
}
static void ConfigureSubgroupSize(const GraphicContext& graphics,
vk::ShaderStageFlagBits vk_stage,
static void ConfigureSubgroupSize(const GraphicContext& graphics, vk::ShaderStageFlagBits vk_stage,
const ShaderRecompiler::IR::Program& program,
vk::PipelineShaderStageRequiredSubgroupSizeCreateInfo& required,
vk::PipelineShaderStageCreateInfo& stage) {
const auto config =
const auto config =
ConfigureShaderSubgroup(ShaderSubgroupCapabilities {graphics}, vk_stage, program);
switch (config.mode) {
case ShaderSubgroupMode::Natural: return;
@@ -456,16 +454,13 @@ static void ConfigureSubgroupSize(const GraphicContext&
}
// NOLINTNEXTLINE(readability-function-cognitive-complexity)
void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline,
const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info,
std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info,
std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0,
uint32_t vs_crc32, uint32_t ps_hash0, uint32_t ps_crc32,
bool ps_active) {
void CreatePipelineInternal(
GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline, const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info, std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info, std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0, uint32_t vs_crc32,
uint32_t ps_hash0, uint32_t ps_crc32, bool ps_active) {
EXIT_IF(ps_active && ps_input_info == nullptr);
vk::ShaderModule vert_shader_module = nullptr;
@@ -511,8 +506,7 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
vert_shader_stage_info.pName = "main";
vert_shader_stage_info.pSpecializationInfo = nullptr;
EXIT_IF(!vs_input_info.stage);
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eVertex,
*vs_input_info.stage.program,
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eVertex, *vs_input_info.stage.program,
vert_subgroup_size, vert_shader_stage_info);
vk::PipelineShaderStageCreateInfo frag_shader_stage_info {};
@@ -527,8 +521,8 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
if (ps_active) {
EXIT_IF(!ps_input_info->stage);
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eFragment,
*ps_input_info->stage.program,
frag_subgroup_size, frag_shader_stage_info);
*ps_input_info->stage.program, frag_subgroup_size,
frag_shader_stage_info);
}
vk::PipelineShaderStageCreateInfo shader_stages[] = {vert_shader_stage_info,
@@ -728,13 +722,13 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
clip_ext.depthClipEnable = static_params.depth_clip_enable ? VK_TRUE : VK_FALSE;
vk::PipelineRasterizationStateCreateInfo rasterizer {};
rasterizer.sType = vk::StructureType::ePipelineRasterizationStateCreateInfo;
rasterizer.sType = vk::StructureType::ePipelineRasterizationStateCreateInfo;
// MoltenVK lacks VK_EXT_depth_clip_enable; omit the depth-clip struct on macOS and accept
// Vulkan's default depth clipping (enabled) instead of the PS5's clamp behavior.
#if defined(__APPLE__)
rasterizer.pNext = nullptr;
rasterizer.pNext = nullptr;
#else
rasterizer.pNext = &clip_ext;
rasterizer.pNext = &clip_ext;
#endif
rasterizer.flags = {};
rasterizer.depthClampEnable = VK_FALSE;
@@ -812,13 +806,13 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
color_write.pColorWriteEnables = color_write_enable;
vk::PipelineColorBlendStateCreateInfo color_blending {};
color_blending.sType = vk::StructureType::ePipelineColorBlendStateCreateInfo;
color_blending.sType = vk::StructureType::ePipelineColorBlendStateCreateInfo;
// MoltenVK lacks VK_EXT_color_write_enable; drop the dynamic color-write struct on macOS
// and rely on each attachment's static colorWriteMask (all channels enabled by default).
#if defined(__APPLE__)
color_blending.pNext = nullptr;
color_blending.pNext = nullptr;
#else
color_blending.pNext = &color_write;
color_blending.pNext = &color_write;
#endif
color_blending.flags = {};
color_blending.logicOpEnable = VK_FALSE;
@@ -838,15 +832,13 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
EXIT_IF(!vs_input_info.stage);
CreateLayout(descriptor_cache, set_layouts, set_layouts_num, push_constant_info,
push_constant_info_num,
*vs_input_info.stage.program, vk::ShaderStageFlagBits::eVertex,
DescriptorCache::Stage::Vertex);
push_constant_info_num, *vs_input_info.stage.program,
vk::ShaderStageFlagBits::eVertex, DescriptorCache::Stage::Vertex);
if (ps_active) {
EXIT_IF(!ps_input_info->stage);
CreateLayout(descriptor_cache, set_layouts, set_layouts_num, push_constant_info,
push_constant_info_num,
*ps_input_info->stage.program, vk::ShaderStageFlagBits::eFragment,
DescriptorCache::Stage::Pixel);
push_constant_info_num, *ps_input_info->stage.program,
vk::ShaderStageFlagBits::eFragment, DescriptorCache::Stage::Pixel);
}
vk::PipelineLayoutCreateInfo pipeline_layout_info {};
@@ -923,32 +915,32 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
dynamic_state.dynamicStateCount = dynamic_states_count;
dynamic_state.pDynamicStates = dynamic_states;
vk::GraphicsPipelineCreateInfo pipeline_info {};
vk::GraphicsPipelineCreateInfo pipeline_info {};
vk::PipelineRenderingCreateInfo rendering_info {};
rendering_info.sType = vk::StructureType::ePipelineRenderingCreateInfo;
rendering_info.colorAttachmentCount = rendering.color_count;
rendering_info.pColorAttachmentFormats = rendering.color_formats.data();
rendering_info.depthAttachmentFormat = rendering.depth_format;
rendering_info.stencilAttachmentFormat = rendering.stencil_format;
pipeline_info.sType = vk::StructureType::eGraphicsPipelineCreateInfo;
pipeline_info.pNext = &rendering_info;
pipeline_info.flags = {};
pipeline_info.stageCount = shader_stage_count;
pipeline_info.pStages = shader_stages;
pipeline_info.pVertexInputState = &vertex_input_info;
pipeline_info.pInputAssemblyState = &input_assembly;
pipeline_info.pTessellationState = nullptr;
pipeline_info.pViewportState = &viewport_state;
pipeline_info.pRasterizationState = &rasterizer;
pipeline_info.pMultisampleState = &multisampling;
pipeline_info.pDepthStencilState = (static_params.with_depth ? &depth_stencil_info : nullptr);
pipeline_info.pColorBlendState = &color_blending;
pipeline_info.pDynamicState = &dynamic_state;
pipeline_info.layout = pipeline.pipeline_layout;
pipeline_info.renderPass = nullptr;
pipeline_info.subpass = 0;
pipeline_info.basePipelineHandle = nullptr;
pipeline_info.basePipelineIndex = -1;
pipeline_info.sType = vk::StructureType::eGraphicsPipelineCreateInfo;
pipeline_info.pNext = &rendering_info;
pipeline_info.flags = {};
pipeline_info.stageCount = shader_stage_count;
pipeline_info.pStages = shader_stages;
pipeline_info.pVertexInputState = &vertex_input_info;
pipeline_info.pInputAssemblyState = &input_assembly;
pipeline_info.pTessellationState = nullptr;
pipeline_info.pViewportState = &viewport_state;
pipeline_info.pRasterizationState = &rasterizer;
pipeline_info.pMultisampleState = &multisampling;
pipeline_info.pDepthStencilState = (static_params.with_depth ? &depth_stencil_info : nullptr);
pipeline_info.pColorBlendState = &color_blending;
pipeline_info.pDynamicState = &dynamic_state;
pipeline_info.layout = pipeline.pipeline_layout;
pipeline_info.renderPass = nullptr;
pipeline_info.subpass = 0;
pipeline_info.basePipelineHandle = nullptr;
pipeline_info.basePipelineIndex = -1;
EXIT_IF(pipeline.pipeline != nullptr);
@@ -1012,8 +1004,7 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
comp_shader_stage_info.pName = "main";
comp_shader_stage_info.pSpecializationInfo = nullptr;
EXIT_IF(!input_info.stage);
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eCompute,
*input_info.stage.program,
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eCompute, *input_info.stage.program,
comp_subgroup_size, comp_shader_stage_info);
vk::DescriptorSetLayout set_layouts[1] = {};
@@ -1024,9 +1015,8 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
EXIT_IF(!input_info.stage);
CreateLayout(descriptor_cache, set_layouts, set_layouts_num, push_constant_info,
push_constant_info_num,
*input_info.stage.program, vk::ShaderStageFlagBits::eCompute,
DescriptorCache::Stage::Compute);
push_constant_info_num, *input_info.stage.program,
vk::ShaderStageFlagBits::eCompute, DescriptorCache::Stage::Compute);
vk::PipelineLayoutCreateInfo pipeline_layout_info {};
pipeline_layout_info.sType = vk::StructureType::ePipelineLayoutCreateInfo;
@@ -10,14 +10,14 @@
#include "graphics/guest_gpu/graphicsRun.h"
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/image/imageInfo.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/pipeline/descriptors.h"
#include "graphics/host_gpu/renderer/image/imageInfo.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/renderer/pipeline/shaderResourceBarrier.h"
#include "graphics/host_gpu/renderer/pipeline/shaderSubgroup.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include "graphics/shader/recompiler/ir/ResourceMaterialization.h"
#include "graphics/shader/recompiler/ir/ShaderIR.h"
@@ -14,8 +14,7 @@ namespace Libs::Graphics {
RenderContext::RenderContext(GraphicContext& graphics)
: m_graphics(graphics), m_render_executor(*this), m_command_scheduler(*this, graphics),
m_descriptor_cache(graphics), m_pipeline_cache(graphics, m_descriptor_cache),
m_sampler_cache(graphics),
m_gpu_resources(graphics, m_command_scheduler) {
m_sampler_cache(graphics), m_gpu_resources(graphics, m_command_scheduler) {
EXIT_NOT_IMPLEMENTED(!Common::Thread::IsMainThread());
}
@@ -27,7 +26,7 @@ RenderContext::~RenderContext() {
void RenderContext::InitializeGpu(VideoOut::VideoOutDriver* video_out) {
EXIT_IF(m_gpu != nullptr);
m_video_out = video_out;
m_gpu = std::make_unique<Gpu>(*this);
m_gpu = std::make_unique<Gpu>(*this);
m_gpu_resources.SetGpu(m_gpu.get());
}
@@ -99,8 +98,7 @@ void RenderContext::TriggerEopEvent(uint32_t context_id) {
registration.eq, static_cast<uintptr_t>(registration.id),
LibKernel::EventQueue::KERNEL_EVFILT_GRAPHICS,
reinterpret_cast<void*>(static_cast<uintptr_t>(context_id)));
if (result == LibKernel::KERNEL_ERROR_EBADF ||
result == LibKernel::KERNEL_ERROR_ENOENT) {
if (result == LibKernel::KERNEL_ERROR_EBADF || result == LibKernel::KERNEL_ERROR_ENOENT) {
DeleteEopEq(registration.eq, registration.id);
continue;
}
+17 -17
View File
@@ -6,12 +6,12 @@
#include "common/common.h"
#include "common/threads.h"
#include "graphics/host_gpu/renderer/cache/bufferCache.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/cache/gpuResourceManager.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "graphics/host_gpu/renderer/cache/samplerCache.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "kernel/eventQueue.h"
#include <memory>
@@ -32,10 +32,10 @@ public:
~RenderContext();
KYTY_CLASS_NO_COPY(RenderContext);
[[nodiscard]] GraphicContext& GetGraphics() const noexcept { return m_graphics; }
void InitializeGpu(VideoOut::VideoOutDriver* video_out);
void ShutdownGpu();
[[nodiscard]] Gpu& GetGpu() const;
[[nodiscard]] GraphicContext& GetGraphics() const noexcept { return m_graphics; }
void InitializeGpu(VideoOut::VideoOutDriver* video_out);
void ShutdownGpu();
[[nodiscard]] Gpu& GetGpu() const;
[[nodiscard]] VideoOut::VideoOutDriver& GetVideoOut() const;
Common::Mutex& GetMutex() { return m_mutex; }
@@ -56,18 +56,18 @@ private:
struct EopEqRegistration {
LibKernel::EventQueue::KernelEqueue eq = LibKernel::EventQueue::KERNEL_EQUEUE_INVALID;
LibKernel::EventQueue::KernelEqueueRef queue;
int id = 0;
int id = 0;
};
GraphicContext& m_graphics;
Common::Mutex m_mutex;
RenderExecutor m_render_executor;
CommandScheduler m_command_scheduler;
DescriptorCache m_descriptor_cache;
PipelineCache m_pipeline_cache;
SamplerCache m_sampler_cache;
GpuResourceManager m_gpu_resources;
std::unique_ptr<Gpu> m_gpu;
GraphicContext& m_graphics;
Common::Mutex m_mutex;
RenderExecutor m_render_executor;
CommandScheduler m_command_scheduler;
DescriptorCache m_descriptor_cache;
PipelineCache m_pipeline_cache;
SamplerCache m_sampler_cache;
GpuResourceManager m_gpu_resources;
std::unique_ptr<Gpu> m_gpu;
VideoOut::VideoOutDriver* m_video_out = nullptr;
Common::Mutex m_eop_mutex;
@@ -12,13 +12,13 @@ namespace Libs::Graphics {
static constexpr uint32_t RENDER_COLOR_ATTACHMENTS_MAX = 8;
struct RenderAttachment {
vk::ImageView image_view = nullptr;
vk::ImageLayout image_layout = vk::ImageLayout::eUndefined;
std::array<uint32_t, 4> clear_value = {};
vk::ImageView image_view = nullptr;
vk::ImageLayout image_layout = vk::ImageLayout::eUndefined;
std::array<uint32_t, 4> clear_value = {};
bool is_clear = false;
bool has_depth = false;
bool depth_clear = false;
bool has_stencil = false;
bool has_depth = false;
bool depth_clear = false;
bool has_stencil = false;
bool stencil_clear = false;
bool operator==(const RenderAttachment&) const = default;
+3 -3
View File
@@ -251,9 +251,9 @@ uint64_t PrepareVideoOutFlip(CommandBuffer& buffer, int handle, int index, int f
int64_t flip_arg) {
for (;;) {
uint64_t request_id = 0;
auto& video_out = buffer.GetContext().GetVideoOut();
const auto result = video_out.SubmitFlipFromGpu(
buffer, handle, index, flip_mode, flip_arg, request_id);
auto& video_out = buffer.GetContext().GetVideoOut();
const auto result =
video_out.SubmitFlipFromGpu(buffer, handle, index, flip_mode, flip_arg, request_id);
if (result == OK) {
EXIT_IF(request_id == 0);
return request_id;
+6 -7
View File
@@ -122,9 +122,9 @@ uint64_t GraphicContext::GetDeviceMemoryUsage() const {
physical_device_properties.deviceType == vk::PhysicalDeviceType::eDiscreteGpu;
uint64_t usage = 0;
for (uint32_t heap = 0; heap < physical_device_memory_properties.memoryHeapCount; heap++) {
const bool device_local = static_cast<bool>(
physical_device_memory_properties.memoryHeaps[heap].flags &
vk::MemoryHeapFlagBits::eDeviceLocal);
const bool device_local =
static_cast<bool>(physical_device_memory_properties.memoryHeaps[heap].flags &
vk::MemoryHeapFlagBits::eDeviceLocal);
if (!discrete || device_local) {
usage += budgets[heap].usage;
}
@@ -144,7 +144,7 @@ uint64_t GraphicContext::GetTotalMemoryBudget() const {
uint64_t local = 0;
uint64_t usage = 0;
for (uint32_t heap = 0; heap < physical_device_memory_properties.memoryHeapCount; heap++) {
const auto& properties = physical_device_memory_properties.memoryHeaps[heap];
const auto& properties = physical_device_memory_properties.memoryHeaps[heap];
const bool device_local =
static_cast<bool>(properties.flags & vk::MemoryHeapFlagBits::eDeviceLocal);
if (device_local) {
@@ -159,9 +159,8 @@ uint64_t GraphicContext::GetTotalMemoryBudget() const {
return budget - std::min<uint64_t>(budget / 8, 1024ull * 1024 * 1024);
}
constexpr uint64_t system_reserve = 8ull * 1024 * 1024 * 1024;
const auto available = budget > usage ? budget - usage : uint64_t {0};
return std::max(local,
available > system_reserve ? available - system_reserve : uint64_t {0});
const auto available = budget > usage ? budget - usage : uint64_t {0};
return std::max(local, available > system_reserve ? available - system_reserve : uint64_t {0});
}
void GraphicContext::CreateBuffer(uint64_t size, VulkanBuffer& buffer) {
+5
View File
@@ -37,6 +37,7 @@ constexpr FormatMapping kFormatMappings[] = {
{Prospero::BufferFormat::k16_16Float, vk::Format::eR16G16Sfloat},
{Prospero::BufferFormat::k11_11_10Float, vk::Format::eB10G11R11UfloatPack32},
{Prospero::BufferFormat::k10_10_10_2UNorm, vk::Format::eA2B10G10R10UnormPack32},
{Prospero::BufferFormat::k10_10_10_2UInt, vk::Format::eA2B10G10R10UintPack32},
{Prospero::BufferFormat::k8_8_8_8UNorm, vk::Format::eR8G8B8A8Unorm},
{Prospero::BufferFormat::k8_8_8_8SNorm, vk::Format::eR8G8B8A8Snorm},
{Prospero::BufferFormat::k8_8_8_8UInt, vk::Format::eR8G8B8A8Uint},
@@ -55,6 +56,10 @@ constexpr FormatMapping kFormatMappings[] = {
{Prospero::BufferFormat::k32_32_32_32UInt, vk::Format::eR32G32B32A32Uint},
{Prospero::BufferFormat::k32_32_32_32SInt, vk::Format::eR32G32B32A32Sint},
{Prospero::BufferFormat::k32_32_32_32Float, vk::Format::eR32G32B32A32Sfloat},
// Narrow-channel sRGB formats are optional in Vulkan. Keep a same-width fallback until
// sampler-aware sRGB emulation is available.
{Prospero::BufferFormat::k8Srgb, vk::Format::eR8Unorm},
{Prospero::BufferFormat::k8_8Srgb, vk::Format::eR8G8Unorm},
{Prospero::BufferFormat::k8_8_8_8Srgb, vk::Format::eR8G8B8A8Srgb},
{Prospero::BufferFormat::k9_9_9_5Float, vk::Format::eE5B9G9R9UfloatPack32},
{Prospero::BufferFormat::k5_6_5UNorm, vk::Format::eB5G6R5UnormPack16},
+7 -7
View File
@@ -20,14 +20,14 @@ public:
~Presenter();
KYTY_CLASS_NO_COPY(Presenter);
[[nodiscard]] Frame& PrepareFrame(CommandBuffer& command, const ImageInfo& info);
[[nodiscard]] Frame& PrepareBlankFrame(uint32_t width, uint32_t height, bool opaque,
CommandBuffer* producer = nullptr);
[[nodiscard]] Frame* PrepareLastFrame();
[[nodiscard]] bool IsGuestPaused() const noexcept;
[[nodiscard]] Frame& PrepareFrame(CommandBuffer& command, const ImageInfo& info);
[[nodiscard]] Frame& PrepareBlankFrame(uint32_t width, uint32_t height, bool opaque,
CommandBuffer* producer = nullptr);
[[nodiscard]] Frame* PrepareLastFrame();
[[nodiscard]] bool IsGuestPaused() const noexcept;
[[nodiscard]] RenderContext& Renderer() const noexcept;
void Present(Frame& frame, bool reuse = false);
void Discard(Frame& frame);
void Present(Frame& frame, bool reuse = false);
void Discard(Frame& frame);
private:
struct Impl;
+37 -41
View File
@@ -69,8 +69,8 @@ enum class FlipRequestSource { Cpu, GpuEop };
struct VideoOutEventState;
struct VideoOutEventRegistration {
EventQueue::KernelEqueue handle = EventQueue::KERNEL_EQUEUE_INVALID;
std::shared_ptr<VideoOutEventState> state;
EventQueue::KernelEqueue handle = EventQueue::KERNEL_EQUEUE_INVALID;
std::shared_ptr<VideoOutEventState> state;
uint64_t generation = 0;
VideoOutEventKind kind = VideoOutEventKind::Flip;
};
@@ -170,13 +170,13 @@ struct BufferAttributeGroup {
struct VideoOutConfig {
Common::Mutex mutex;
Common::CondVar vblank_cond;
std::shared_ptr<VideoOutEventState> events = std::make_shared<VideoOutEventState>();
uint32_t width = 0;
uint32_t height = 0;
uint64_t generation = 0;
bool opened = false;
bool closing = false;
int flip_rate = 0;
std::shared_ptr<VideoOutEventState> events = std::make_shared<VideoOutEventState>();
uint32_t width = 0;
uint32_t height = 0;
uint64_t generation = 0;
bool opened = false;
bool closing = false;
int flip_rate = 0;
uint64_t output_mode = VIDEO_OUT_OUTPUT_MODE_DEFAULT;
float gamma = 1.0f;
VideoOutFlipStatus flip_status;
@@ -250,8 +250,8 @@ public:
VideoOutConfig* Get(int handle, uint64_t& generation);
bool IsOpened(int handle);
void Init(uint32_t width, uint32_t height);
FlipQueue& GetFlipQueue() { return m_flip_queue; }
void Init(uint32_t width, uint32_t height);
FlipQueue& GetFlipQueue() { return m_flip_queue; }
Graphics::RenderContext& Renderer() const noexcept { return m_renderer; }
void VblankBegin();
@@ -259,12 +259,12 @@ public:
void PresentThread(std::stop_token token);
private:
Common::Mutex m_mutex;
VideoOutConfig m_video_out_ctx[VIDEO_OUT_NUM_MAX];
Common::Mutex m_mutex;
VideoOutConfig m_video_out_ctx[VIDEO_OUT_NUM_MAX];
Graphics::RenderContext& m_renderer;
Graphics::Presenter& m_presenter;
FlipQueue m_flip_queue;
std::jthread m_present_thread;
Graphics::Presenter& m_presenter;
FlipQueue m_flip_queue;
std::jthread m_present_thread;
};
static std::unique_ptr<VideoOutDriver> g_video_out_driver;
@@ -279,7 +279,7 @@ static uintptr_t VideoOutEventId(VideoOutEventKind kind) {
}
static VideoOutEventQueues& VideoOutEventQueuesFor(VideoOutEventState& state,
VideoOutEventKind kind) {
VideoOutEventKind kind) {
switch (kind) {
case VideoOutEventKind::Flip: return state.flip;
case VideoOutEventKind::Vblank: return state.vblank;
@@ -359,9 +359,9 @@ static void TriggerVideoOutEvents(VideoOutConfig& video_out, VideoOutEventKind k
if (!registration || registration->generation != video_out.generation) {
continue;
}
const auto result = EventQueue::KernelTriggerEvent(
registration->handle, VideoOutEventId(kind), EventQueue::KERNEL_EVFILT_VIDEO_OUT,
trigger_data);
const auto result =
EventQueue::KernelTriggerEvent(registration->handle, VideoOutEventId(kind),
EventQueue::KERNEL_EVFILT_VIDEO_OUT, trigger_data);
EXIT_NOT_IMPLEMENTED(result != OK && result != LibKernel::KERNEL_ERROR_EBADF &&
result != LibKernel::KERNEL_ERROR_ENOENT);
}
@@ -372,9 +372,8 @@ static void DeleteVideoOutEvents(const VideoOutEventQueues& queues, VideoOutEven
if (!registration) {
continue;
}
const auto result =
EventQueue::KernelDeleteEvent(registration->handle, VideoOutEventId(kind),
EventQueue::KERNEL_EVFILT_VIDEO_OUT);
const auto result = EventQueue::KernelDeleteEvent(
registration->handle, VideoOutEventId(kind), EventQueue::KERNEL_EVFILT_VIDEO_OUT);
EXIT_NOT_IMPLEMENTED(result != OK && result != LibKernel::KERNEL_ERROR_EBADF &&
result != LibKernel::KERNEL_ERROR_ENOENT);
}
@@ -383,7 +382,7 @@ static void DeleteVideoOutEvents(const VideoOutEventQueues& queues, VideoOutEven
static int RegisterVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoOutEventKind kind,
void* udata) {
uint64_t generation = 0;
auto* video_out = DriverState().Get(handle, generation);
auto* video_out = DriverState().Get(handle, generation);
if (video_out == nullptr) {
return VIDEO_OUT_ERROR_INVALID_HANDLE;
}
@@ -425,27 +424,25 @@ static int RegisterVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoO
bool add_queue = false;
{
Common::LockGuard event_lock(event_state->mutex);
const auto existing = std::find_if(queues.begin(), queues.end(), [&](const auto& candidate) {
return candidate->handle == eq && candidate->generation == generation;
});
const auto existing =
std::find_if(queues.begin(), queues.end(), [&](const auto& candidate) {
return candidate->handle == eq && candidate->generation == generation;
});
if (existing != queues.end()) {
registration = *existing;
} else {
registration = std::make_shared<VideoOutEventRegistration>(
VideoOutEventRegistration {.handle = eq,
.state = event_state,
.generation = generation,
.kind = kind});
registration = std::make_shared<VideoOutEventRegistration>(VideoOutEventRegistration {
.handle = eq, .state = event_state, .generation = generation, .kind = kind});
queues.push_back(registration);
add_queue = true;
}
}
event.filter.data = registration.get();
event.filter.owner = registration;
const int result = EventQueue::KernelAddEvent(eq, event);
const int result = EventQueue::KernelAddEvent(eq, event);
if (result != OK && add_queue) {
Common::LockGuard event_lock(event_state->mutex);
const auto added = std::find(queues.begin(), queues.end(), registration);
const auto added = std::find(queues.begin(), queues.end(), registration);
if (added != queues.end()) {
queues.erase(added);
}
@@ -455,7 +452,7 @@ static int RegisterVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoO
static int DeleteVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoOutEventKind kind) {
uint64_t generation = 0;
auto* video_out = DriverState().Get(handle, generation);
auto* video_out = DriverState().Get(handle, generation);
if (video_out == nullptr) {
return VIDEO_OUT_ERROR_INVALID_HANDLE;
}
@@ -814,8 +811,8 @@ void VideoOutDriver::Impl::PresentThread(std::stop_token token) {
m_presenter.Present(*frame, true);
}
const auto frame_end = Common::Timer::QueryPerformanceCounter();
total_wait += static_cast<int64_t>(period) -
static_cast<int64_t>(frame_end - frame_begin);
total_wait +=
static_cast<int64_t>(period) - static_cast<int64_t>(frame_end - frame_begin);
continue;
}
@@ -841,8 +838,7 @@ void VideoOutDriver::Impl::PresentThread(std::stop_token token) {
VblankEnd();
const auto frame_end = Common::Timer::QueryPerformanceCounter();
total_wait += static_cast<int64_t>(period) -
static_cast<int64_t>(frame_end - frame_begin);
total_wait += static_cast<int64_t>(period) - static_cast<int64_t>(frame_end - frame_begin);
}
}
@@ -1000,8 +996,8 @@ void FlipQueue::Prepare(uint64_t request_id, Graphics::CommandBuffer& buffer) {
}
Graphics::Presenter::Frame* frame = nullptr;
if (special) {
frame = &m_presenter.PrepareBlankFrame(width, height,
index == VIDEO_OUT_BUFFER_INDEX_BLACK, &buffer);
frame = &m_presenter.PrepareBlankFrame(width, height, index == VIDEO_OUT_BUFFER_INDEX_BLACK,
&buffer);
} else {
frame = &m_presenter.PrepareFrame(buffer, source_info);
}
+7 -7
View File
@@ -32,13 +32,13 @@ public:
~VideoOutDriver();
KYTY_CLASS_NO_COPY(VideoOutDriver);
int SubmitFlipFromGpu(Graphics::CommandBuffer& buffer, int handle, int index, int flip_mode,
int64_t flip_arg, uint64_t& request_id);
void PrepareFlip(uint64_t request_id, Graphics::CommandBuffer& buffer);
void CompleteFlip(uint64_t request_id);
void SubmitFlipPreparation(uint64_t request_id);
void WaitForSubmitSlot();
void WaitFlipDone(int handle, int index);
int SubmitFlipFromGpu(Graphics::CommandBuffer& buffer, int handle, int index, int flip_mode,
int64_t flip_arg, uint64_t& request_id);
void PrepareFlip(uint64_t request_id, Graphics::CommandBuffer& buffer);
void CompleteFlip(uint64_t request_id);
void SubmitFlipPreparation(uint64_t request_id);
void WaitForSubmitSlot();
void WaitFlipDone(int handle, int index);
[[nodiscard]] Impl& State() noexcept;
+61 -83
View File
@@ -61,7 +61,7 @@ namespace Libs::Graphics {
struct Presenter::Frame {
VulkanImage image;
std::unique_ptr<CommandBuffer> present_commands;
bool busy = false;
bool busy = false;
bool reusing_last = false;
void Configure(GraphicContext& graphics, vk::Extent2D extent, vk::Format format);
@@ -155,7 +155,7 @@ public:
EXIT("last submitted frame is not available for reuse\n");
}
m_free.erase(free);
m_last_frame = nullptr;
m_last_frame = nullptr;
frame->busy = true;
frame->reusing_last = true;
m_mutex.Unlock();
@@ -197,30 +197,27 @@ private:
}
}
WindowContext& m_window;
Common::Mutex m_mutex;
Common::CondVar m_available;
WindowContext& m_window;
Common::Mutex m_mutex;
Common::CondVar m_available;
std::vector<std::unique_ptr<Presenter::Frame>> m_frames;
std::deque<Presenter::Frame*> m_free;
Presenter::Frame* m_last_frame = nullptr;
vk::Format m_format = vk::Format::eUndefined;
vk::Format m_format = vk::Format::eUndefined;
};
void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent,
vk::Format format) {
void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent, vk::Format format) {
if (extent.width == 0 || extent.height == 0 || format == vk::Format::eUndefined) {
EXIT("unsupported prepared frame, extent=%ux%u format=%d\n", extent.width, extent.height,
static_cast<int>(format));
}
const auto features = graphics.GetFormatProperties(format).optimalTilingFeatures;
const auto required = vk::FormatFeatureFlagBits::eBlitSrc |
vk::FormatFeatureFlagBits::eSampledImageFilterLinear |
vk::FormatFeatureFlagBits::eTransferSrc |
vk::FormatFeatureFlagBits::eTransferDst;
const auto required =
vk::FormatFeatureFlagBits::eBlitSrc | vk::FormatFeatureFlagBits::eSampledImageFilterLinear |
vk::FormatFeatureFlagBits::eTransferSrc | vk::FormatFeatureFlagBits::eTransferDst;
if ((features & required) != required) {
EXIT("prepared presentation format lacks optimal blit support: format=%d features=0x%x\n",
static_cast<int>(format),
static_cast<vk::FormatFeatureFlags::MaskType>(features));
static_cast<int>(format), static_cast<vk::FormatFeatureFlags::MaskType>(features));
}
auto& dst = image;
@@ -234,11 +231,11 @@ void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent,
dst.memory = {};
}
dst.extent = {extent.width, extent.height, 1};
dst.format = format;
dst.layers = 1;
dst.mip_levels = 1;
dst.state = {};
dst.extent = {extent.width, extent.height, 1};
dst.format = format;
dst.layers = 1;
dst.mip_levels = 1;
dst.state = {};
dst.subresource_states.clear();
dst.memory.property = vk::MemoryPropertyFlagBits::eDeviceLocal;
@@ -262,13 +259,12 @@ void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent,
void Presenter::Frame::Transit(vk::CommandBuffer command, vk::ImageLayout layout,
vk::AccessFlags2 access) {
const auto stage = access == vk::AccessFlagBits2::eTransferRead ||
access == vk::AccessFlagBits2::eTransferWrite
? vk::PipelineStageFlagBits2::eTransfer
: vk::PipelineStageFlagBits2::eAllCommands;
const auto stage = access == vk::AccessFlagBits2::eTransferRead ||
access == vk::AccessFlagBits2::eTransferWrite
? vk::PipelineStageFlagBits2::eTransfer
: vk::PipelineStageFlagBits2::eAllCommands;
constexpr auto writes = vk::AccessFlagBits2::eTransferWrite |
vk::AccessFlagBits2::eShaderWrite |
vk::AccessFlagBits2::eMemoryWrite;
vk::AccessFlagBits2::eShaderWrite | vk::AccessFlagBits2::eMemoryWrite;
if (image.state.layout == layout && image.state.access_mask == access &&
!static_cast<bool>(image.state.access_mask & writes)) {
return;
@@ -299,35 +295,27 @@ void Presenter::Frame::Transit(vk::CommandBuffer command, vk::ImageLayout layout
void Presenter::Frame::CopyFrom(CommandBuffer& command_buffer, Image& source) {
command_buffer.EndRendering();
auto command = command_buffer.Handle();
source.Transit(vk::ImageLayout::eTransferSrcOptimal,
vk::AccessFlagBits2::eTransferRead, {}, command);
Transit(command, vk::ImageLayout::eTransferDstOptimal,
vk::AccessFlagBits2::eTransferWrite);
source.Transit(vk::ImageLayout::eTransferSrcOptimal, vk::AccessFlagBits2::eTransferRead, {},
command);
Transit(command, vk::ImageLayout::eTransferDstOptimal, vk::AccessFlagBits2::eTransferWrite);
vk::ImageCopy copy {};
copy.srcSubresource = {vk::ImageAspectFlagBits::eColor, 0, 0,
source.backing.layers};
copy.srcSubresource = {vk::ImageAspectFlagBits::eColor, 0, 0, source.backing.layers};
copy.dstSubresource = {vk::ImageAspectFlagBits::eColor, 0, 0, image.layers};
copy.extent = {std::min(source.backing.extent.width, image.extent.width),
std::min(source.backing.extent.height, image.extent.height), 1};
copy.extent = {std::min(source.backing.extent.width, image.extent.width),
std::min(source.backing.extent.height, image.extent.height), 1};
EXIT_IF(copy.srcSubresource.layerCount != copy.dstSubresource.layerCount);
command.copyImage(source.backing.image, vk::ImageLayout::eTransferSrcOptimal,
image.image, vk::ImageLayout::eTransferDstOptimal, copy);
Transit(command, vk::ImageLayout::eTransferSrcOptimal,
vk::AccessFlagBits2::eTransferRead);
command.copyImage(source.backing.image, vk::ImageLayout::eTransferSrcOptimal, image.image,
vk::ImageLayout::eTransferDstOptimal, copy);
Transit(command, vk::ImageLayout::eTransferSrcOptimal, vk::AccessFlagBits2::eTransferRead);
}
void Presenter::Frame::Clear(CommandBuffer& command_buffer,
const vk::ClearColorValue& color) {
void Presenter::Frame::Clear(CommandBuffer& command_buffer, const vk::ClearColorValue& color) {
command_buffer.EndRendering();
auto command = command_buffer.Handle();
Transit(command, vk::ImageLayout::eTransferDstOptimal,
vk::AccessFlagBits2::eTransferWrite);
const vk::ImageSubresourceRange range {
vk::ImageAspectFlagBits::eColor, 0, 1, 0, 1};
command.clearColorImage(image.image, vk::ImageLayout::eTransferDstOptimal, &color, 1,
&range);
Transit(command, vk::ImageLayout::eTransferSrcOptimal,
vk::AccessFlagBits2::eTransferRead);
Transit(command, vk::ImageLayout::eTransferDstOptimal, vk::AccessFlagBits2::eTransferWrite);
const vk::ImageSubresourceRange range {vk::ImageAspectFlagBits::eColor, 0, 1, 0, 1};
command.clearColorImage(image.image, vk::ImageLayout::eTransferDstOptimal, &color, 1, &range);
Transit(command, vk::ImageLayout::eTransferSrcOptimal, vk::AccessFlagBits2::eTransferRead);
}
class Swapchain final {
@@ -338,8 +326,8 @@ public:
~Swapchain();
KYTY_CLASS_NO_COPY(Swapchain);
void Create();
void Recreate(bool surface_lost = false);
void Create();
void Recreate(bool surface_lost = false);
[[nodiscard]] Status AcquireNextImage();
void RecordPresentCommands(CommandBuffer& command, VulkanImage& source);
void Submit(CommandBuffer& command);
@@ -395,17 +383,17 @@ struct Presenter::Impl {
desc.view_info.usage = vk::ImageUsageFlagBits::eTransferSrc;
desc.type = TextureCache::BindingType::VideoOut;
auto& cache = renderer.GetTextureCache();
auto& image = cache.GetImage(cache.FindImage(desc));
auto& cache = renderer.GetTextureCache();
auto& image = cache.GetImage(cache.FindImage(desc));
image.usage.video_out = true;
return image;
}
RenderContext& renderer;
WindowContext& window;
Swapchain swapchain;
RenderContext& renderer;
WindowContext& window;
Swapchain swapchain;
CommandScheduler present_scheduler;
FramePool frames;
FramePool frames;
};
void Swapchain::Create() {
@@ -441,25 +429,20 @@ void Swapchain::Create() {
? vk::CompositeAlphaFlagBitsKHR::eOpaque
: vk::CompositeAlphaFlagBitsKHR::eInherit;
vk::SurfaceFormatKHR format {vk::Format::eR8G8B8A8Unorm,
vk::ColorSpaceKHR::eSrgbNonlinear};
if (surface.formats.size() != 1 ||
surface.formats.front().format != vk::Format::eUndefined) {
vk::SurfaceFormatKHR format {vk::Format::eR8G8B8A8Unorm, vk::ColorSpaceKHR::eSrgbNonlinear};
if (surface.formats.size() != 1 || surface.formats.front().format != vk::Format::eUndefined) {
const auto it = std::find_if(surface.formats.begin(), surface.formats.end(),
[](const vk::SurfaceFormatKHR& candidate) {
return candidate.format ==
vk::Format::eB8G8R8A8Unorm ||
candidate.format ==
vk::Format::eR8G8B8A8Unorm;
return candidate.format == vk::Format::eB8G8R8A8Unorm ||
candidate.format == vk::Format::eR8G8B8A8Unorm;
});
if (it == surface.formats.end()) {
EXIT("no supported UNORM swapchain format\n");
}
format = *it;
}
m_format = format.format;
const auto swapchain_features =
graphics.GetFormatProperties(m_format).optimalTilingFeatures;
m_format = format.format;
const auto swapchain_features = graphics.GetFormatProperties(m_format).optimalTilingFeatures;
if (!static_cast<bool>(swapchain_features & vk::FormatFeatureFlagBits::eBlitDst)) {
EXIT("swapchain format cannot be a blit destination: format=%d\n",
static_cast<int>(m_format));
@@ -503,9 +486,8 @@ void Swapchain::Create() {
view.subresourceRange.baseMipLevel = 0;
view.subresourceRange.layerCount = 1;
view.subresourceRange.levelCount = 1;
RequireVulkanSuccess(
graphics.device.createImageView(&view, nullptr, &m_image_views[i]),
"vkCreateImageView");
RequireVulkanSuccess(graphics.device.createImageView(&view, nullptr, &m_image_views[i]),
"vkCreateImageView");
EXIT_IF(m_image_views[i] == nullptr);
}
@@ -600,7 +582,7 @@ void Swapchain::Recreate(bool surface_lost) {
Swapchain::Status Swapchain::AcquireNextImage() {
EXIT_IF(m_handle == nullptr || m_frame_index >= m_image_acquired.size());
m_image_index = static_cast<uint32_t>(-1);
m_image_index = static_cast<uint32_t>(-1);
const auto result = m_window.graphic_ctx.device.acquireNextImageKHR(
m_handle, std::numeric_limits<uint64_t>::max(), m_image_acquired[m_frame_index], nullptr,
&m_image_index);
@@ -683,10 +665,9 @@ void Swapchain::RecordPresentCommands(CommandBuffer& command, VulkanImage& sourc
to_present.subresourceRange.levelCount = 1;
to_present.subresourceRange.baseArrayLayer = 0;
to_present.subresourceRange.layerCount = 1;
vk_command.pipelineBarrier(vk::PipelineStageFlagBits::eAllCommands,
vk::PipelineStageFlagBits::eAllCommands,
vk::DependencyFlagBits::eByRegion, 0,
nullptr, 0, nullptr, 1, &to_present);
vk_command.pipelineBarrier(
vk::PipelineStageFlagBits::eAllCommands, vk::PipelineStageFlagBits::eAllCommands,
vk::DependencyFlagBits::eByRegion, 0, nullptr, 0, nullptr, 1, &to_present);
command.End();
}
@@ -700,7 +681,7 @@ void Swapchain::Submit(CommandBuffer& command) {
Swapchain::Status Swapchain::Present() {
EXIT_IF(m_image_index >= m_render_complete.size());
const auto ready = m_render_complete[m_image_index];
const auto ready = m_render_complete[m_image_index];
vk::PresentInfoKHR present {};
present.sType = vk::StructureType::ePresentInfoKHR;
present.swapchainCount = 1;
@@ -738,7 +719,7 @@ Presenter::~Presenter() = default;
Presenter::Frame& Presenter::PrepareFrame(CommandBuffer& buffer, const ImageInfo& info) {
KYTY_PROFILER_FUNCTION();
EXIT_IF(buffer.IsInvalid());
auto* frame = m_impl->frames.Acquire();
auto* frame = m_impl->frames.Acquire();
Common::LockGuard render_lock(m_impl->renderer.GetMutex());
auto& image = m_impl->ResolveSurface(info);
if (image.backing.format == vk::Format::eUndefined) {
@@ -752,14 +733,13 @@ Presenter::Frame& Presenter::PrepareFrame(CommandBuffer& buffer, const ImageInfo
default: break;
}
frame->Configure(m_impl->window.graphic_ctx,
{image.backing.extent.width, image.backing.extent.height},
frame_format);
{image.backing.extent.width, image.backing.extent.height}, frame_format);
frame->CopyFrom(buffer, image);
return *frame;
}
Presenter::Frame& Presenter::PrepareBlankFrame(uint32_t width, uint32_t height, bool opaque,
CommandBuffer* producer) {
CommandBuffer* producer) {
KYTY_PROFILER_FUNCTION();
auto format = m_impl->frames.GetFormat();
auto* frame = m_impl->frames.Acquire();
@@ -772,8 +752,7 @@ Presenter::Frame& Presenter::PrepareBlankFrame(uint32_t width, uint32_t height,
frame->Clear(*producer, clear);
} else {
if (frame->present_commands == nullptr) {
frame->present_commands =
std::make_unique<CommandBuffer>(m_impl->present_scheduler);
frame->present_commands = std::make_unique<CommandBuffer>(m_impl->present_scheduler);
}
auto& command = *frame->present_commands;
command.WaitForFenceAndReset();
@@ -830,8 +809,7 @@ void Presenter::Present(Frame& frame, bool reuse) {
continue;
}
if (frame.present_commands == nullptr) {
frame.present_commands =
std::make_unique<CommandBuffer>(m_impl->present_scheduler);
frame.present_commands = std::make_unique<CommandBuffer>(m_impl->present_scheduler);
}
{
Common::LockGuard render_lock(m_impl->renderer.GetMutex());
@@ -32,11 +32,11 @@
#include "graphics/host_gpu/vma.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include "graphics/presentation/presenter.h"
#include "kernel/memory.h"
#include "graphics/presentation/renderDoc.h"
#include "graphics/presentation/videoOut.h"
#include "graphics/presentation/window.h"
#include "graphics/presentation/window/windowInternal.h"
#include "kernel/memory.h"
#include "libs/controller.h"
#include "loader/systemContent.h"
@@ -287,6 +287,10 @@ static void VulkanFindPhysicalDevice(vk::Instance instance, vk::SurfaceKHR surfa
LOGF("shaderStorageImageReadWithoutFormat is not supported\n");
skip_device = true;
}
if (features12.shaderStorageImageArrayNonUniformIndexing != VK_TRUE) {
LOGF("shaderStorageImageArrayNonUniformIndexing is not supported\n");
skip_device = true;
}
if (device_features2.features.shaderImageGatherExtended != VK_TRUE) {
LOGF("shaderImageGatherExtended is not supported\n");
@@ -475,9 +479,9 @@ static void VulkanInitSubgroupSizeControl(vk::PhysicalDevice physical_device,
}
static vk::Device VulkanCreateDevice(vk::PhysicalDevice physical_device, const VulkanExtensions& r,
uint32_t queue_family,
uint32_t queue_family,
const std::vector<const char*>& device_extensions,
GraphicContext& graphics) {
GraphicContext& graphics) {
EXIT_IF(physical_device == nullptr);
EXIT_IF(queue_family == static_cast<uint32_t>(-1));
@@ -514,6 +518,7 @@ static vk::Device VulkanCreateDevice(vk::PhysicalDevice physical_device, const V
features12.sType = vk::StructureType::ePhysicalDeviceVulkan12Features;
features12.pNext = &depth_clip_control;
features12.samplerMirrorClampToEdge = VK_TRUE;
features12.shaderStorageImageArrayNonUniformIndexing = VK_TRUE;
vk::PhysicalDeviceSubgroupSizeControlFeatures subgroup_size_control {};
subgroup_size_control.sType = vk::StructureType::ePhysicalDeviceSubgroupSizeControlFeatures;
@@ -551,19 +556,19 @@ static vk::Device VulkanCreateDevice(vk::PhysicalDevice physical_device, const V
features12.timelineSemaphore = VK_TRUE;
vk::PhysicalDeviceFeatures device_features {};
device_features.fragmentStoresAndAtomics = VK_TRUE;
device_features.samplerAnisotropy = VK_TRUE;
device_features.robustBufferAccess = VK_TRUE;
device_features.fragmentStoresAndAtomics = VK_TRUE;
device_features.samplerAnisotropy = VK_TRUE;
device_features.robustBufferAccess = VK_TRUE;
#if !defined(__APPLE__)
device_features.depthBounds = VK_TRUE; // unsupported by MoltenVK
device_features.depthBounds = VK_TRUE; // unsupported by MoltenVK
#endif
device_features.shaderStorageImageWriteWithoutFormat = VK_TRUE;
device_features.shaderStorageImageReadWithoutFormat = VK_TRUE;
device_features.shaderImageGatherExtended = VK_TRUE;
device_features.independentBlend = VK_TRUE;
device_features.tessellationShader = VK_TRUE;
device_features.sampleRateShading = VK_TRUE;
graphics.sample_rate_shading_enabled = true;
device_features.sampleRateShading = VK_TRUE;
graphics.sample_rate_shading_enabled = true;
device_features.vertexPipelineStoresAndAtomics =
supported_features2.features.vertexPipelineStoresAndAtomics;
@@ -909,10 +914,9 @@ void WindowContext::CreateVulkan() {
}
surface = native_surface;
std::vector<const char*> device_extensions = {VK_KHR_SWAPCHAIN_EXTENSION_NAME,
VK_EXT_DEPTH_CLIP_CONTROL_EXTENSION_NAME,
VK_KHR_PUSH_DESCRIPTOR_EXTENSION_NAME,
"VK_KHR_maintenance1"};
std::vector<const char*> device_extensions = {
VK_KHR_SWAPCHAIN_EXTENSION_NAME, VK_EXT_DEPTH_CLIP_CONTROL_EXTENSION_NAME,
VK_KHR_PUSH_DESCRIPTOR_EXTENSION_NAME, "VK_KHR_maintenance1"};
#if defined(__APPLE__)
// MoltenVK lacks VK_EXT_depth_clip_enable and VK_EXT_color_write_enable; the renderer
@@ -932,8 +936,8 @@ void WindowContext::CreateVulkan() {
uint32_t queue_family = static_cast<uint32_t>(-1);
VulkanFindPhysicalDevice(graphic_ctx.instance, surface, device_extensions,
surface_capabilities, graphic_ctx.physical_device, queue_family);
VulkanFindPhysicalDevice(graphic_ctx.instance, surface, device_extensions, surface_capabilities,
graphic_ctx.physical_device, queue_family);
if (graphic_ctx.physical_device == nullptr) {
EXIT("Could not find suitable device");
@@ -949,9 +953,8 @@ void WindowContext::CreateVulkan() {
auto available_extensions = EnumerateVulkan<vk::ExtensionProperties>(
"vkEnumerateDeviceExtensionProperties",
[&](uint32_t* count, vk::ExtensionProperties* values) {
return graphic_ctx.physical_device.enumerateDeviceExtensionProperties(nullptr,
count,
values);
return graphic_ctx.physical_device.enumerateDeviceExtensionProperties(
nullptr, count, values);
});
if (HasExtension(available_extensions, VK_EXT_MEMORY_BUDGET_EXTENSION_NAME)) {
@@ -985,7 +988,7 @@ void WindowContext::CreateVulkan() {
render_context = std::make_unique<RenderContext>(graphic_ctx);
LibKernel::Memory::InstallGpuResources(&render_context->GetGpuResources());
presenter = std::make_unique<Presenter>(*this);
presenter = std::make_unique<Presenter>(*this);
RenderDocSetActiveWindow(graphic_ctx.instance, window);
}
+23 -25
View File
@@ -1,7 +1,5 @@
#include "graphics/presentation/window.h"
#include <cstdlib>
#include "SDL.h"
#include "SDL_error.h"
#include "SDL_events.h"
@@ -40,6 +38,7 @@
#include <algorithm>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <memory>
#include <string>
@@ -59,7 +58,7 @@
namespace Libs::Graphics {
constexpr int KEYBOARD_CONTROLLER_ID = -1000;
constexpr int KEYBOARD_CONTROLLER_ID = -1000;
struct EventKeyboard {
bool down;
@@ -251,9 +250,7 @@ static void GameEventKeyboard(WindowLoopState& game, const EventKeyboard& key) {
if (key.down) {
switch (key.key_code) {
case SDLK_ESCAPE: game.need_exit = true; break;
case SDLK_SPACE:
SetPause(game, !game.paused.load(std::memory_order_acquire));
break;
case SDLK_SPACE: SetPause(game, !game.paused.load(std::memory_order_acquire)); break;
case SDLK_F1:
if (!key.repeat) {
RenderDocRequestCapture();
@@ -390,7 +387,9 @@ void WindowContext::Resize(uint32_t new_width, uint32_t new_height) {
void WindowContext::ProcessWindowEvent(const SDL_WindowEvent& event) {
const auto& window_event = event;
switch (window_event.event) {
case SDL_WINDOWEVENT_SHOWN: LOGF("Window %" PRIu32 " shown\n", window_event.windowID); break;
case SDL_WINDOWEVENT_SHOWN:
LOGF("Window %" PRIu32 " shown\n", window_event.windowID);
break;
case SDL_WINDOWEVENT_HIDDEN:
LOGF("Window %" PRIu32 " hidden\n", window_event.windowID);
@@ -401,13 +400,13 @@ void WindowContext::ProcessWindowEvent(const SDL_WindowEvent& event) {
break;
case SDL_WINDOWEVENT_MOVED:
LOGF("Window %" PRIu32 " moved to %" PRId32 ",%" PRId32 "\n",
window_event.windowID, window_event.data1, window_event.data2);
LOGF("Window %" PRIu32 " moved to %" PRId32 ",%" PRId32 "\n", window_event.windowID,
window_event.data1, window_event.data2);
break;
case SDL_WINDOWEVENT_RESIZED:
LOGF("Window %" PRIu32 " resized to %" PRId32 "x%" PRId32 "\n",
window_event.windowID, window_event.data1, window_event.data2);
LOGF("Window %" PRIu32 " resized to %" PRId32 "x%" PRId32 "\n", window_event.windowID,
window_event.data1, window_event.data2);
LOGF("m: %d\n", static_cast<int>(SDL_ThreadID()));
Resize(window_event.data1, window_event.data2);
@@ -807,9 +806,8 @@ static void WindowCreate(WindowContext& context) {
window_flags |= static_cast<uint32_t>(SDL_WINDOW_BORDERLESS);
}
#endif
context.window =
SDL_CreateWindow(KYTY_SDL_WINDOW_CAPTION, KYTY_SDL_WINDOWPOS_CENTERED,
KYTY_SDL_WINDOWPOS_CENTERED, width, height, window_flags);
context.window = SDL_CreateWindow(KYTY_SDL_WINDOW_CAPTION, KYTY_SDL_WINDOWPOS_CENTERED,
KYTY_SDL_WINDOWPOS_CENTERED, width, height, window_flags);
context.window_hidden = true;
@@ -832,7 +830,7 @@ Presenter& WindowInit(uint32_t width, uint32_t height) {
WindowCreate(*window);
window->CreateVulkan();
auto& presenter = *window->presenter;
g_window = std::move(window);
g_window = std::move(window);
return presenter;
}
@@ -934,9 +932,9 @@ void WindowContext::UpdateTitle() {
Loader::SystemContentParamSfoGetString("TITLE_ID", title_id, sizeof(title_id));
static bool has_app_ver =
Loader::SystemContentParamSfoGetString("APP_VER", app_ver, sizeof(app_ver));
static uint64_t fps_start = Common::Timer::QueryPerformanceCounter();
static uint64_t frame_num = 0;
static uint64_t fps_frames = 0;
static uint64_t fps_start = Common::Timer::QueryPerformanceCounter();
static uint64_t frame_num = 0;
static uint64_t fps_frames = 0;
static double current_fps = 0.0;
const auto now = Common::Timer::QueryPerformanceCounter();
@@ -946,15 +944,15 @@ void WindowContext::UpdateTitle() {
if (now - fps_start >= frequency) {
current_fps = static_cast<double>(fps_frames) * static_cast<double>(frequency) /
static_cast<double>(now - fps_start);
fps_start = now;
fps_frames = 0;
fps_start = now;
fps_frames = 0;
}
auto fps = fmt::format("{}{}{}{}{}{}[{}] [{}], frame: {}, fps: {:f}", (has_title ? title : ""),
(has_title ? ", " : ""), (has_title_id ? title_id : ""),
(has_title_id ? ", " : ""), (has_app_ver ? app_ver : ""),
(has_app_ver ? " " : ""), device_name, processor_name,
frame_num, current_fps);
auto fps =
fmt::format("{}{}{}{}{}{}[{}] [{}], frame: {}, fps: {:f}", (has_title ? title : ""),
(has_title ? ", " : ""), (has_title_id ? title_id : ""),
(has_title_id ? ", " : ""), (has_app_ver ? app_ver : ""),
(has_app_ver ? " " : ""), device_name, processor_name, frame_num, current_fps);
#if defined(__APPLE__)
// AppKit traps on title changes off the main thread; fire-and-forget keeps present pacing.
@@ -28,8 +28,8 @@ struct SurfaceCapabilities {
};
struct WindowLoopState {
SDL_Event event {};
bool need_exit = false;
SDL_Event event {};
bool need_exit = false;
std::atomic_bool paused = false;
};
@@ -38,14 +38,13 @@ struct WindowContext {
~WindowContext();
KYTY_CLASS_NO_COPY(WindowContext);
[[nodiscard]] static vk::PhysicalDeviceVulkan13Features
RequiredVulkan13Features() noexcept;
void CreateVulkan();
void RecreateSurface();
void RefreshSurfaceCapabilities();
void UpdateIcon();
void UpdateTitle();
void Resize(uint32_t width, uint32_t height);
[[nodiscard]] static vk::PhysicalDeviceVulkan13Features RequiredVulkan13Features() noexcept;
void CreateVulkan();
void RecreateSurface();
void RefreshSurfaceCapabilities();
void UpdateIcon();
void UpdateTitle();
void Resize(uint32_t width, uint32_t height);
void ProcessWindowEvent(const SDL_WindowEvent& event);
void ProcessDisplayEvent(const SDL_DisplayEvent& event);
void ProcessEvent(double time_seconds);
@@ -59,14 +58,14 @@ struct WindowContext {
void DrainMainThreadTasks();
#endif
GraphicContext graphic_ctx;
SDL_Window* window = nullptr;
bool window_hidden = true;
vk::SurfaceKHR surface = nullptr;
SurfaceCapabilities surface_capabilities;
GraphicContext graphic_ctx;
SDL_Window* window = nullptr;
bool window_hidden = true;
vk::SurfaceKHR surface = nullptr;
SurfaceCapabilities surface_capabilities;
std::unique_ptr<RenderContext> render_context;
std::unique_ptr<Presenter> presenter;
WindowLoopState loop;
std::unique_ptr<Presenter> presenter;
WindowLoopState loop;
char device_name[VK_MAX_PHYSICAL_DEVICE_NAME_SIZE] = {0};
char processor_name[64] = {0};
@@ -76,7 +75,7 @@ struct WindowContext {
#if defined(__APPLE__)
Common::Mutex main_task_mutex;
Common::CondVar main_task_done;
std::vector<std::function<void()>> main_tasks; // guarded by main_task_mutex
std::vector<std::function<void()>> main_tasks; // guarded by main_task_mutex
uint64_t main_tasks_queued = 0; // guarded by main_task_mutex
uint64_t main_tasks_run = 0; // guarded by main_task_mutex
#endif
@@ -2,15 +2,16 @@
#include "common/assert.h"
#include "common/logging/log.h"
#include "graphics/shader/recompiler/cfg/ShaderCFG.h"
#include "graphics/shader/recompiler/decompiler/ShaderDecoder.h"
#include "graphics/shader/recompiler/emitter/SpirvEmitter.h"
#include "graphics/shader/recompiler/ir/BindingLayout.h"
#include "graphics/shader/recompiler/ir/ReadLaneElimination.h"
#include "graphics/shader/recompiler/ir/ResourceMaterialization.h"
#include "graphics/shader/recompiler/ir/ResourceTracking.h"
#include "graphics/shader/recompiler/ir/ScalarProvenance.h"
#include "graphics/shader/recompiler/cfg/ShaderCFG.h"
#include "graphics/shader/recompiler/decompiler/ShaderDecoder.h"
#include "graphics/shader/recompiler/ir/ShaderIR.h"
#include "graphics/shader/recompiler/ir/ShaderInfoCollection.h"
#include "graphics/shader/recompiler/emitter/SpirvEmitter.h"
#include "graphics/shader/recompiler/ir/SrtPatcher.h"
#include "graphics/shader/recompiler/ir/SrtWalker.h"
@@ -838,6 +839,11 @@ bool TryRecompile(std::span<const uint32_t> code, const CompileOptions& options,
if (!IR::AllocateBindings(ir, layout_options, error)) {
return false;
}
const auto read_lane_stats = IR::EliminateReadLane(ir);
if (read_lane_stats.rewritten_reads != 0) {
LOGF("%s read-lane elimination: reads=%" PRIu32 " shadow_writes=%" PRIu32 "\n",
GetDumpLabel(options), read_lane_stats.rewritten_reads, read_lane_stats.shadow_writes);
}
std::string ir_dump;
if (options.dump_ir) {
ir_dump = MakeIrDump(cfg, ir);
@@ -44,7 +44,7 @@ struct CompileResult {
};
bool TryRecompile(std::span<const uint32_t> code, const CompileOptions& options,
CompileResult& result, std::string* error);
CompileResult& result, std::string* error);
} // namespace Libs::Graphics::ShaderRecompiler
+248 -19
View File
@@ -871,19 +871,19 @@ std::vector<uint32_t> DominatedBlocks(const Graph& graph, uint32_t header,
return blocks;
}
uint32_t AppendSyntheticMergeBlock(Graph& graph, uint32_t old_merge) {
const auto* merge = graph.FindBlock(old_merge);
uint32_t AppendSyntheticBranchBlock(Graph& graph, uint32_t target) {
const auto* target_block = graph.FindBlock(target);
BasicBlock block;
block.id = static_cast<uint32_t>(graph.blocks.size());
block.start_pc = merge != nullptr ? merge->start_pc : 0u;
block.start_pc = target_block != nullptr ? target_block->start_pc : 0u;
block.end_pc = block.start_pc;
block.inst_begin = merge != nullptr ? merge->inst_begin : 0u;
block.inst_begin = target_block != nullptr ? target_block->inst_begin : 0u;
block.inst_end = block.inst_begin;
block.successors = {old_merge};
block.successors = {target};
block.terminator.kind = TerminatorKind::Branch;
block.terminator.condition = BranchCondition::Always;
block.terminator.true_block = old_merge;
block.terminator.true_block = target;
graph.blocks.push_back(std::move(block));
return graph.blocks.back().id;
}
@@ -897,15 +897,50 @@ bool IsSyntheticMergeForwarder(const Graph& graph, uint32_t block_id, uint32_t m
block->terminator.true_block == merge;
}
bool IsInsideLoopConstruct(const Graph& graph, const NaturalLoop& loop, uint32_t block_id) {
return block_id != UINT32_MAX && block_id != loop.merge && block_id != loop.continue_block &&
graph.Dominates(loop.header, block_id) &&
(loop.merge == UINT32_MAX || !graph.Dominates(loop.merge, block_id));
const NaturalLoop* FindInnermostContainingLoop(const Graph& graph, uint32_t block_id) {
const NaturalLoop* innermost = nullptr;
for (const auto& loop: graph.natural_loops) {
if (Contains(loop.body_blocks, block_id) &&
(innermost == nullptr || loop.body_blocks.size() < innermost->body_blocks.size())) {
innermost = &loop;
}
}
return innermost;
}
bool SelectionMergeLeavesContainingLoop(const Graph& graph, uint32_t header, uint32_t merge) {
bool IsInsideLoopConstruct(const Graph& graph, const NaturalLoop& loop, uint32_t block_id) {
return block_id != UINT32_MAX && block_id != loop.merge && block_id != loop.continue_block &&
graph.Dominates(loop.header, block_id) && !graph.Dominates(loop.merge, block_id);
}
bool IsInnermostLoopControlConditional(const Graph& graph, const BasicBlock& block) {
if (block.terminator.kind != TerminatorKind::ConditionalBranch) {
return false;
}
const auto* loop = FindInnermostContainingLoop(graph, block.id);
if (loop == nullptr || loop->merge == UINT32_MAX || loop->continue_block == UINT32_MAX) {
return false;
}
const auto true_target = block.terminator.true_block;
const auto false_target = block.terminator.false_block;
if (block.id == loop->continue_block) {
const auto is_repeat_target = [&](uint32_t target) {
return target == loop->header || target == loop->merge;
};
return is_repeat_target(true_target) && is_repeat_target(false_target);
}
const auto is_control_target = [&](uint32_t target) {
return target == loop->merge || target == loop->continue_block;
};
return (is_control_target(true_target) &&
(is_control_target(false_target) ||
IsInsideLoopConstruct(graph, *loop, false_target))) ||
(is_control_target(false_target) && IsInsideLoopConstruct(graph, *loop, true_target));
}
bool MergeLeavesContainingLoop(const Graph& graph, uint32_t header, uint32_t merge) {
for (const auto& loop: graph.natural_loops) {
if (IsInsideLoopConstruct(graph, loop, header) &&
if (loop.header != header && IsInsideLoopConstruct(graph, loop, header) &&
!IsInsideLoopConstruct(graph, loop, merge)) {
return true;
}
@@ -913,6 +948,80 @@ bool SelectionMergeLeavesContainingLoop(const Graph& graph, uint32_t header, uin
return false;
}
bool CanonicalizeNaturalLoops(Graph& graph, std::string* error) {
const auto rewrite_budget = graph.blocks.size() * 2u + 16u;
for (size_t rewrite = 0; rewrite < rewrite_budget; rewrite++) {
bool changed = false;
for (const auto& loop: graph.natural_loops) {
std::vector<uint32_t> latches;
for (const auto& edge: graph.back_edges) {
if (edge.to == loop.header) {
AddUnique(latches, edge.from);
}
}
if (latches.size() <= 1u) {
continue;
}
const auto continue_block = AppendSyntheticBranchBlock(graph, loop.header);
for (auto latch: latches) {
auto* block = graph.FindBlock(latch);
if (block != nullptr) {
ReplaceValue(block->successors, loop.header, continue_block);
ReplaceTerminatorTarget(block->terminator, loop.header, continue_block);
}
}
RebuildPredecessors(graph);
RecomputeAnalyses(graph);
changed = true;
break;
}
if (changed) {
continue;
}
for (const auto& loop: graph.natural_loops) {
const auto* header = graph.FindBlock(loop.header);
const auto is_loop_control_target = [&](uint32_t target) {
return target == loop.merge || target == loop.continue_block;
};
if (header == nullptr || header->terminator.kind != TerminatorKind::ConditionalBranch ||
is_loop_control_target(header->terminator.true_block) ||
is_loop_control_target(header->terminator.false_block) ||
!Contains(loop.body_blocks, header->terminator.true_block) ||
!Contains(loop.body_blocks, header->terminator.false_block)) {
continue;
}
const auto old_header = loop.header;
const auto predecessors = header->predecessors;
const auto new_header = AppendSyntheticBranchBlock(graph, old_header);
for (auto pred: predecessors) {
auto* block = graph.FindBlock(pred);
if (block != nullptr) {
ReplaceValue(block->successors, old_header, new_header);
ReplaceTerminatorTarget(block->terminator, old_header, new_header);
}
}
if (graph.entry_block == old_header) {
graph.entry_block = new_header;
}
MoveBlockBefore(graph, new_header, old_header);
RebuildPredecessors(graph);
RecomputeAnalyses(graph);
changed = true;
break;
}
if (!changed) {
return true;
}
}
SetFailure(graph, FailureKind::StructuredControlFlow, graph.entry_block,
"CFG loop canonicalization exceeded rewrite budget", error);
return false;
}
bool SplitSharedMergeBlock(Graph& graph, uint32_t merge,
const std::vector<uint32_t>& construct_blocks,
bool force_split = false) {
@@ -948,7 +1057,7 @@ bool SplitSharedMergeBlock(Graph& graph, uint32_t merge,
return false;
}
const auto synthetic_merge = AppendSyntheticMergeBlock(graph, merge);
const auto synthetic_merge = AppendSyntheticBranchBlock(graph, merge);
auto* synthetic_block = graph.FindBlock(synthetic_merge);
if (synthetic_block != nullptr) {
synthetic_block->predecessors = predecessors_to_split;
@@ -980,14 +1089,111 @@ bool SplitSharedMergeBlock(Graph& graph, uint32_t merge,
bool SplitOneLoopMerge(Graph& graph) {
const auto& loops = graph.natural_loops;
for (const auto& loop: loops) {
if (SplitSharedMergeBlock(graph, loop.merge, loop.body_blocks)) {
const auto construct_blocks = DominatedBlocks(graph, loop.header, loop.merge);
const auto force_split = MergeLeavesContainingLoop(graph, loop.header, loop.merge);
if (SplitSharedMergeBlock(graph, loop.merge, construct_blocks, force_split)) {
return true;
}
}
return false;
}
bool SplitOneSelectionMerge(Graph& graph) {
std::vector<uint32_t> SelectionRegion(const Graph& graph, const BasicBlock& header,
uint32_t merge) {
std::vector<uint32_t> region;
std::vector<uint32_t> pending = {header.terminator.true_block,
header.terminator.false_block};
while (!pending.empty()) {
const auto block_id = pending.back();
pending.pop_back();
if (block_id == merge || Contains(region, block_id)) {
continue;
}
const auto* block = graph.FindBlock(block_id);
if (block == nullptr) {
continue;
}
AddUnique(region, block_id);
pending.insert(pending.end(), block->successors.begin(), block->successors.end());
}
SortUnique(region);
return region;
}
bool DuplicateSelectionRegion(Graph& graph, uint32_t header_id, uint32_t merge,
const std::vector<uint32_t>& region, uint32_t block_budget) {
std::vector<uint32_t> cloned_blocks;
for (auto block_id: region) {
if (!graph.Dominates(header_id, block_id)) {
cloned_blocks.push_back(block_id);
}
}
if (cloned_blocks.empty() || graph.FindBlock(header_id) == nullptr || header_id >= merge ||
graph.blocks.size() + cloned_blocks.size() + 1u > block_budget) {
return false;
}
const auto first_clone = static_cast<uint32_t>(graph.blocks.size());
std::map<uint32_t, uint32_t> clones;
for (uint32_t i = 0; i < cloned_blocks.size(); i++) {
clones.emplace(cloned_blocks[i], first_clone + i);
}
for (auto block_id: cloned_blocks) {
BasicBlock clone = *graph.FindBlock(block_id);
clone.id = clones.at(block_id);
clone.predecessors.clear();
clone.dominators.clear();
clone.post_dominators.clear();
graph.blocks.push_back(std::move(clone));
}
const auto remap_block = [&](BasicBlock& block) {
const auto remap_target = [&](uint32_t& target) {
if (const auto it = clones.find(target); it != clones.end()) {
target = it->second;
}
};
for (auto& successor: block.successors) {
remap_target(successor);
}
remap_target(block.terminator.true_block);
remap_target(block.terminator.false_block);
remap_target(block.terminator.merge_block);
remap_target(block.terminator.continue_block);
for (auto& target: block.terminator.indirect_targets) {
remap_target(target);
}
};
for (auto block_id: region) {
const auto owned_id = clones.contains(block_id) ? clones.at(block_id) : block_id;
remap_block(*graph.FindBlock(owned_id));
}
const auto private_merge = AppendSyntheticBranchBlock(graph, merge);
auto& header = *graph.FindBlock(header_id);
remap_block(header);
for (auto block_id: region) {
const auto owned_id = clones.contains(block_id) ? clones.at(block_id) : block_id;
auto* block = graph.FindBlock(owned_id);
if (block != nullptr) {
ReplaceValue(block->successors, merge, private_merge);
ReplaceTerminatorTarget(block->terminator, merge, private_merge);
}
}
ReplaceValue(header.successors, merge, private_merge);
ReplaceTerminatorTarget(header.terminator, merge, private_merge);
for (uint32_t i = 0; i <= cloned_blocks.size(); i++) {
MoveBlockBefore(graph, first_clone + i, merge + i);
}
RebuildPredecessors(graph);
RecomputeAnalyses(graph);
return true;
}
bool SplitOneSelectionMerge(Graph& graph, uint32_t block_budget) {
std::vector<uint32_t> loop_headers;
loop_headers.reserve(graph.natural_loops.size());
for (const auto& loop: graph.natural_loops) {
@@ -1001,11 +1207,26 @@ bool SplitOneSelectionMerge(Graph& graph) {
Contains(loop_headers, block_id)) {
continue;
}
if (IsInnermostLoopControlConditional(graph, *block)) {
continue;
}
const auto merge = graph.FindNearestCommonPostDominator(block->terminator.true_block,
block->terminator.false_block);
if (merge == UINT32_MAX || graph.FindBlock(merge) == nullptr) {
continue;
}
const auto region = SelectionRegion(graph, *block, merge);
if (std::any_of(region.begin(), region.end(),
[&](uint32_t member) { return !graph.Dominates(block_id, member); })) {
if (graph.natural_loops.empty() &&
DuplicateSelectionRegion(graph, block_id, merge, region, block_budget)) {
return true;
}
continue;
}
const auto construct_blocks = DominatedBlocks(graph, block_id, merge);
const auto force_split = SelectionMergeLeavesContainingLoop(graph, block_id, merge);
const auto force_split = MergeLeavesContainingLoop(graph, block_id, merge);
if (SplitSharedMergeBlock(graph, merge, construct_blocks, force_split)) {
return true;
}
@@ -1015,10 +1236,12 @@ bool SplitOneSelectionMerge(Graph& graph) {
bool SplitSharedMergeBlocks(Graph& graph, std::string* error) {
const auto original_block_count = static_cast<uint32_t>(graph.blocks.size());
const auto split_budget =
std::max<uint32_t>(16u, std::min<uint32_t>(128u, original_block_count));
const auto split_budget = std::max<uint32_t>(
16u, std::min<uint32_t>(128u, original_block_count * 4u));
const auto block_budget = std::max<uint32_t>(
32u, std::min<uint32_t>(512u, original_block_count * 8u));
for (uint32_t splits = 0; splits < split_budget; splits++) {
if (!SplitOneLoopMerge(graph) && !SplitOneSelectionMerge(graph)) {
if (!SplitOneLoopMerge(graph) && !SplitOneSelectionMerge(graph, block_budget)) {
return true;
}
RebuildPredecessors(graph);
@@ -1353,6 +1576,9 @@ bool Structurize(Graph& graph, std::string* error) {
return false;
}
if (!CanonicalizeNaturalLoops(graph, error)) {
return false;
}
if (!SplitSharedMergeBlocks(graph, error)) {
return false;
}
@@ -1395,6 +1621,9 @@ bool Structurize(Graph& graph, std::string* error) {
block.terminator.loop_header) {
continue;
}
if (IsInnermostLoopControlConditional(graph, block)) {
continue;
}
const auto merge = graph.FindNearestCommonPostDominator(block.terminator.true_block,
block.terminator.false_block);
@@ -35,9 +35,9 @@ constexpr ImageDimension DecodeImageDimension(uint32_t dim) {
case 2u: return ImageDimension::Dim3D;
case 3u: return ImageDimension::Dim2DArray;
case 4u: return ImageDimension::Dim1DArray;
case 5u:
case 7u: return ImageDimension::Dim2DArray;
case 6u: return ImageDimension::Dim2D;
case 5u: return ImageDimension::Dim2DArray;
case 6u: return ImageDimension::Dim2DMsaa;
case 7u: return ImageDimension::Dim2DMsaaArray;
default: return ImageDimension::Unknown;
}
}
@@ -46,8 +46,10 @@ constexpr uint32_t ImageCoordComponents(ImageDimension dimension) {
switch (dimension) {
case ImageDimension::Dim1D: return 1u;
case ImageDimension::Dim1DArray: return 2u;
case ImageDimension::Dim2DMsaa:
case ImageDimension::Dim3D:
case ImageDimension::Dim2DArray: return 3u;
case ImageDimension::Dim2DMsaaArray: return 4u;
default: return 2u;
}
}
@@ -196,10 +196,10 @@ bool DecodeSopk(uint32_t pc, std::span<const uint32_t> code, uint32_t word_index
case Opcode::SMovkI32: return DecodeScalarDestination(sdst, pc, inst.dst, error);
case Opcode::SWaitcnt: {
const uint32_t waitcnt = word & 0xffffu;
inst.dst.kind = OperandKind::Null;
inst.src0.signed_val = static_cast<int32_t>(waitcnt);
inst.src0.value = waitcnt;
inst.src_count = 1;
inst.dst.kind = OperandKind::Null;
inst.src0.signed_val = static_cast<int32_t>(waitcnt);
inst.src0.value = waitcnt;
inst.src_count = 1;
return true;
}
case Opcode::SSetregB32:
@@ -266,10 +266,10 @@ bool DecodeSopp(uint32_t pc, std::span<const uint32_t> code, uint32_t word_index
inst.src0.value = simm;
inst.src0.signed_val = static_cast<int16_t>(simm);
inst.src_count = (inst.opcode == Opcode::SNop || inst.opcode == Opcode::SWaitcnt ||
inst.opcode == Opcode::SSleep || inst.opcode == Opcode::SSendmsg ||
inst.opcode == Opcode::STtraceData || inst.opcode == Opcode::SInstPrefetch)
? 1
: 0;
inst.opcode == Opcode::SSleep || inst.opcode == Opcode::SSendmsg ||
inst.opcode == Opcode::STtraceData || inst.opcode == Opcode::SInstPrefetch)
? 1
: 0;
inst.branch_offset = static_cast<int32_t>(static_cast<int16_t>(simm)) * 4;
inst.branch_target = pc + 4u + static_cast<uint32_t>(inst.branch_offset);
SetRawWords(inst, code, word_index, 1);
@@ -194,6 +194,8 @@ const char* ImageDimensionToString(ImageDimension dimension) {
case ImageDimension::Dim2D: return "2d";
case ImageDimension::Dim3D: return "3d";
case ImageDimension::Dim2DArray: return "2d_array";
case ImageDimension::Dim2DMsaa: return "2d_msaa";
case ImageDimension::Dim2DMsaaArray: return "2d_msaa_array";
default: return "unknown";
}
}
@@ -220,9 +222,9 @@ bool DecodeScalarSource(uint32_t code, uint32_t pc, Operand& operand, std::strin
}
if (code >= 240u && code <= 247u) {
constexpr float values[] = {0.5f, -0.5f, 1.0f, -1.0f, 2.0f, -2.0f, 4.0f, -4.0f};
operand.kind = OperandKind::FloatInlineConstant;
operand.float_val = values[code - 240u];
operand.value = FloatBits(operand.float_val);
operand.kind = OperandKind::FloatInlineConstant;
operand.float_val = values[code - 240u];
operand.value = FloatBits(operand.float_val);
return true;
}
if (code >= 256u && code <= 511u) {
@@ -285,7 +287,7 @@ bool DecodeVectorGpr(uint32_t reg, Operand& operand, std::string* error) {
SetError(error, "VGPR index is out of range");
return false;
}
operand = {};
operand = {};
operand.kind = OperandKind::Vgpr;
operand.reg = reg;
return true;
@@ -575,6 +575,8 @@ enum class ImageDimension : uint32_t {
Dim2D,
Dim3D,
Dim2DArray,
Dim2DMsaa,
Dim2DMsaaArray,
};
constexpr uint32_t MaxInstructionRawWords = 5u;
@@ -30,10 +30,16 @@ bool ImageBinding(const IR::ImageResource& image, IR::DescriptorBindingKind& kin
kind = integer ? Kind::SampledUint1DArray : Kind::Sampled1DArray;
return true;
case Dim::Dim2D: kind = integer ? Kind::SampledUint2D : Kind::Sampled2D; return true;
case Dim::Dim2DMsaa:
kind = integer ? Kind::SampledUint2DMsaa : Kind::Sampled2DMsaa;
return true;
case Dim::Dim3D: kind = integer ? Kind::SampledUint3D : Kind::Sampled3D; return true;
case Dim::Dim2DArray:
kind = integer ? Kind::SampledUint2DArray : Kind::Sampled2DArray;
return true;
case Dim::Dim2DMsaaArray:
kind = integer ? Kind::SampledUint2DMsaaArray : Kind::Sampled2DMsaaArray;
return true;
case Dim::Unknown: return false;
}
}
@@ -51,6 +57,8 @@ bool ImageBinding(const IR::ImageResource& image, IR::DescriptorBindingKind& kin
case Dim::Dim2DArray:
kind = uint_image ? Kind::StorageUint2DArray : Kind::Storage2DArray;
return true;
case Dim::Dim2DMsaa:
case Dim::Dim2DMsaaArray: return false;
case Dim::Unknown: return false;
}
return false;
@@ -229,8 +237,14 @@ bool ValidateNativeProgram(const IR::Program& program, std::string* error) {
if (!ImageBinding(program.info.images[i], kind)) {
return Fail(error, "native shader plan has an invalid image class");
}
const auto bindings = program.info.images[i].NumBindings();
if (bindings == 0 || bindings > IR::ImageResource::MaxMipLevels) {
return Fail(error, "native shader plan has an invalid image descriptor count");
}
present[static_cast<size_t>(kind)] = true;
expected[static_cast<size_t>(kind)].push_back(i);
for (uint32_t binding = 0; binding < bindings; binding++) {
expected[static_cast<size_t>(kind)].push_back(i);
}
}
if (!program.info.samplers.empty()) {
Expect(Kind::Samplers, Dense(program.info.samplers.size()));
@@ -308,6 +322,11 @@ bool ValidateNativeProgram(const IR::Program& program, std::string* error) {
inst.memory.image_dimension)) {
return Fail(error, "image instruction has an invalid dense resource");
}
if (inst.op == IR::Opcode::ImageStore &&
((program.info.images[inst.memory.resource].mip_mode ==
IR::ImageMipMode::DynamicStorage) != inst.memory.image_has_mip)) {
return Fail(error, "storage image mip mode does not match the instruction");
}
const bool address =
inst.op == IR::Opcode::SLoadDword || memory == IR::ResourceKind::Flat ||
memory == IR::ResourceKind::Global || memory == IR::ResourceKind::Scratch;
@@ -12,9 +12,9 @@ namespace Libs::Graphics::ShaderRecompiler::Spirv {
bool ProgramRequiresExactSubgroupSize(const IR::Program& program);
bool EmitProgram(const IR::Program& program, const IR::ResourceSnapshot& resources,
const ShaderVertexInputInfo* vertex_input_info,
const ShaderPixelInputInfo* pixel_input_info,
const ShaderComputeInputInfo* compute_input_info, std::vector<uint32_t>& spirv,
const ShaderVertexInputInfo* vertex_input_info,
const ShaderPixelInputInfo* pixel_input_info,
const ShaderComputeInputInfo* compute_input_info, std::vector<uint32_t>& spirv,
std::string* error);
} // namespace Libs::Graphics::ShaderRecompiler::Spirv
@@ -129,7 +129,7 @@ uint32_t MaxCollectedVectorRegisterEnd(const std::vector<RegisterBinding>& regis
}
void CollectMoveRelSourceRegisters(const IR::Program& program,
std::vector<RegisterBinding>& registers) {
std::vector<RegisterBinding>& registers) {
const auto max_vector_end = MaxCollectedVectorRegisterEnd(registers);
for (const auto& block: program.blocks) {
for (const auto& inst: block.instructions) {
@@ -276,8 +276,7 @@ void CopyProgramInputsAndOutputs(EmitterState& state, const IR::Program& program
if (HasOutput(state.outputs, output.kind, output.index)) {
continue;
}
state.outputs.push_back(
{output.kind, output.index, output.location, 0, output.debug_name});
state.outputs.push_back({output.kind, output.index, output.location, 0, output.debug_name});
}
}
@@ -561,12 +560,21 @@ uint32_t DescriptorElementPointer(EmitterState& state, uint32_t result_ptr_type,
uint32_t variable_id, uint32_t array_index,
IR::DescriptorBindingKind kind, uint32_t resource,
const char* variable_name) {
return DescriptorElementPointerId(state, result_ptr_type, variable_id,
ConstantU32(state, array_index), kind, resource,
variable_name);
}
uint32_t DescriptorElementPointerId(EmitterState& state, uint32_t result_ptr_type,
uint32_t variable_id, uint32_t array_index_id,
IR::DescriptorBindingKind kind, uint32_t resource,
const char* variable_name) {
if (variable_id == 0) {
ExitDescriptorBindingFailure(state, kind, resource, variable_name);
}
const auto pointer = state.builder.AllocateId();
state.builder.AddFunction(
{OpAccessChain, result_ptr_type, pointer, variable_id, ConstantU32(state, array_index)});
{OpAccessChain, result_ptr_type, pointer, variable_id, array_index_id});
return pointer;
}
@@ -576,6 +584,8 @@ ImageViewKind ImageViewKindFromDimension(Decoder::ImageDimension dimension) {
case Decoder::ImageDimension::Dim1DArray: return ImageViewKind::Dim1DArray;
case Decoder::ImageDimension::Dim2DArray: return ImageViewKind::Dim2DArray;
case Decoder::ImageDimension::Dim3D: return ImageViewKind::Dim3D;
case Decoder::ImageDimension::Dim2DMsaa: return ImageViewKind::Dim2DMsaa;
case Decoder::ImageDimension::Dim2DMsaaArray: return ImageViewKind::Dim2DMsaaArray;
default: return ImageViewKind::Dim2D;
}
}
@@ -601,7 +611,9 @@ uint32_t ImageViewCoordinateComponents(ImageViewKind view) {
case ImageViewKind::Dim1DArray:
case ImageViewKind::Dim2D: return 2u;
case ImageViewKind::Dim2DArray:
case ImageViewKind::Dim2DMsaaArray:
case ImageViewKind::Dim3D: return 3u;
case ImageViewKind::Dim2DMsaa: return 2u;
default: return 0u;
}
}
@@ -611,7 +623,9 @@ uint32_t ImageViewSpatialComponents(ImageViewKind view) {
case ImageViewKind::Dim1D:
case ImageViewKind::Dim1DArray: return 1u;
case ImageViewKind::Dim2D:
case ImageViewKind::Dim2DArray: return 2u;
case ImageViewKind::Dim2DArray:
case ImageViewKind::Dim2DMsaa:
case ImageViewKind::Dim2DMsaaArray: return 2u;
case ImageViewKind::Dim3D: return 3u;
default: return 0u;
}
@@ -663,8 +677,7 @@ uint32_t LoadSampledImageDescriptor(EmitterState& state, const IR::MemoryInfo& m
uint32_t LoadSamplerDescriptor(EmitterState& state, uint32_t sampler, uint32_t use_pc) {
(void)use_pc;
const auto binding =
ResourceForDescriptor(state, IR::DescriptorBindingKind::Samplers, sampler);
const auto binding = ResourceForDescriptor(state, IR::DescriptorBindingKind::Samplers, sampler);
const auto pointer = DescriptorElementPointer(
state, state.ptr_uniform_sampler, state.sampler_variable, binding.array_index,
IR::DescriptorBindingKind::Samplers, sampler, "sampler descriptor array was not emitted");
@@ -24,7 +24,7 @@ uint32_t EmitExportVec4F32(EmitterState& state, const IR::Instruction& inst) {
const auto raw = EmitValueLoad(state, inst.src[pair_index]);
const auto unpacked = state.builder.AllocateId();
state.builder.AddFunction({OpExtInst, state.vec2_float_type, unpacked,
state.glsl_std450, GlslUnpackHalf2x16, raw});
state.glsl_std450, GlslUnpackHalf2x16, raw});
for (uint32_t lane = 0; lane < 2u; lane++) {
const auto component = pair_index * 2u + lane;
if (((inst.export_info.en >> component) & 1u) == 0) {
@@ -36,8 +36,8 @@ uint32_t EmitExportVec4F32(EmitterState& state, const IR::Instruction& inst) {
}
}
const auto vec = state.builder.AllocateId();
state.builder.AddFunction({OpCompositeConstruct, state.vec4_float_type, vec,
components[0], components[1], components[2], components[3]});
state.builder.AddFunction({OpCompositeConstruct, state.vec4_float_type, vec, components[0],
components[1], components[2], components[3]});
return vec;
}
@@ -50,7 +50,59 @@ uint32_t EmitExportVec4F32(EmitterState& state, const IR::Instruction& inst) {
return vec;
}
uint32_t ApplyMrtExportMapping(EmitterState& state, const IR::Instruction& inst, uint32_t value) {
uint32_t EmitExportComponentU32(EmitterState& state, const IR::Instruction& inst,
uint32_t component) {
const bool enabled = ((inst.export_info.en >> component) & 1u) != 0;
if (!enabled || component >= inst.src_count || component >= 4u) {
return ConstantU32(state, component == 3u ? 1u : 0u);
}
return EmitValueLoad(state, inst.src[component]);
}
uint32_t EmitExportVec4U32(EmitterState& state, const IR::Instruction& inst) {
uint32_t components[4] = {
ConstantU32(state, 0u),
ConstantU32(state, 0u),
ConstantU32(state, 0u),
ConstantU32(state, 1u),
};
if (inst.export_info.compr) {
for (uint32_t pair_index = 0; pair_index < 2u && pair_index < inst.src_count;
pair_index++) {
const auto raw = EmitValueLoad(state, inst.src[pair_index]);
for (uint32_t lane = 0; lane < 2u; lane++) {
const auto component = pair_index * 2u + lane;
if (((inst.export_info.en >> component) & 1u) == 0) {
continue;
}
components[component] = state.builder.AllocateId();
state.builder.AddFunction(
{OpBitFieldUExtract, state.uint_type, components[component], raw,
ConstantU32(state, lane * 16u), ConstantU32(state, 16u)});
}
}
} else {
for (uint32_t component = 0; component < 4u; component++) {
components[component] = EmitExportComponentU32(state, inst, component);
}
}
const auto vec = state.builder.AllocateId();
state.builder.AddFunction({OpCompositeConstruct, state.vec4_uint_type, vec, components[0],
components[1], components[2], components[3]});
return vec;
}
static bool MrtUsesUintOutput(const EmitterState& state, const IR::Instruction& inst) {
return inst.export_info.kind == IR::ExportTargetKind::Mrt &&
state.pixel_input_info != nullptr &&
inst.export_info.index < std::size(state.pixel_input_info->target_output_mode) &&
state.pixel_input_info->target_output_mode[inst.export_info.index] == 7u;
}
uint32_t ApplyMrtExportMapping(EmitterState& state, const IR::Instruction& inst, uint32_t value,
uint32_t vector_type) {
if (inst.export_info.kind != IR::ExportTargetKind::Mrt || state.pixel_input_info == nullptr ||
inst.export_info.index >= state.pixel_input_info->target_export_mapping.size()) {
return value;
@@ -62,8 +114,8 @@ uint32_t ApplyMrtExportMapping(EmitterState& state, const IR::Instruction& inst,
}
const auto mapped = state.builder.AllocateId();
state.builder.AddFunction({OpVectorShuffle, state.vec4_float_type, mapped, value, value,
mapping.Map(0), mapping.Map(1), mapping.Map(2), mapping.Map(3)});
state.builder.AddFunction({OpVectorShuffle, vector_type, mapped, value, value, mapping.Map(0),
mapping.Map(1), mapping.Map(2), mapping.Map(3)});
return mapped;
}
@@ -89,7 +141,7 @@ void EmitMrtZExport(EmitterState& state, const IR::Instruction& inst) {
const auto ptr = state.builder.AllocateId();
state.builder.AddFunction({OpBitcast, state.int_type, mask, raw});
state.builder.AddFunction({OpAccessChain, state.ptr_output_int, ptr,
state.sample_mask_variable, ConstantU32(state, 0)});
state.sample_mask_variable, ConstantU32(state, 0)});
state.builder.AddFunction({OpStore, ptr, mask});
}
}
@@ -114,11 +166,15 @@ void EmitExport(EmitterState& state, const IR::Instruction& inst) {
return;
}
const auto value = ApplyMrtExportMapping(state, inst, EmitExportVec4F32(state, inst));
const auto uint_output = MrtUsesUintOutput(state, inst);
const auto vector_type = uint_output ? state.vec4_uint_type : state.vec4_float_type;
const auto value = ApplyMrtExportMapping(
state, inst, uint_output ? EmitExportVec4U32(state, inst) : EmitExportVec4F32(state, inst),
vector_type);
if (inst.export_info.kind == IR::ExportTargetKind::Position) {
const auto pointer = state.builder.AllocateId();
state.builder.AddFunction({OpAccessChain, state.ptr_output_vec4_float, pointer, variable,
ConstantU32(state, 0)});
state.builder.AddFunction(
{OpAccessChain, state.ptr_output_vec4_float, pointer, variable, ConstantU32(state, 0)});
state.builder.AddFunction({OpStore, pointer, value});
return;
}
@@ -102,7 +102,7 @@ uint32_t EmitWqmLaneU32(EmitterState& state, uint32_t src) {
state.builder.AddFunction(
{OpINotEqual, state.bool_type, non_zero, masked, ConstantU32(state, 0)});
state.builder.AddFunction({OpSelect, state.uint_type, expanded, non_zero,
ConstantU32(state, mask), ConstantU32(state, 0)});
ConstantU32(state, mask), ConstantU32(state, 0)});
state.builder.AddFunction({OpBitwiseOr, state.uint_type, combined, ret, expanded});
ret = combined;
}
@@ -122,8 +122,8 @@ void EmitWqmB64(EmitterState& state, const IR::Instruction& inst) {
}
const auto ballot = state.builder.AllocateId();
state.builder.AddFunction({OpGroupNonUniformBallot, state.vec4_uint_type, ballot,
ConstantU32(state, ScopeSubgroup),
EmitLaneMaskOperandActiveBool(state, inst.src[0])});
ConstantU32(state, ScopeSubgroup),
EmitLaneMaskOperandActiveBool(state, inst.src[0])});
const auto low = state.builder.AllocateId();
const auto high = state.builder.AllocateId();
state.builder.AddFunction({OpCompositeExtract, state.uint_type, low, ballot, 0});
@@ -150,8 +150,8 @@ void EmitWqmB64(EmitterState& state, const IR::Instruction& inst) {
EmitPerInvocationMask(state, inst.dst, active);
} else {
const auto result = state.builder.AllocateId();
state.builder.AddFunction({OpSelect, state.uint_type, result, active,
ConstantU32(state, 1), ConstantU32(state, 0)});
state.builder.AddFunction({OpSelect, state.uint_type, result, active, ConstantU32(state, 1),
ConstantU32(state, 0)});
EmitStoreU32(state, inst.dst, result);
EmitStoreU32(state, OffsetRegisterOperand(inst.dst, 1), ConstantU32(state, 0));
}
@@ -205,8 +205,7 @@ void EmitSaveexecB32(EmitterState& state, const IR::Instruction& inst) {
const auto cond = state.builder.AllocateId();
const auto scc = state.builder.AllocateId();
state.builder.AddFunction(
{OpINotEqual, state.bool_type, cond, new_low, ConstantU32(state, 0)});
state.builder.AddFunction({OpINotEqual, state.bool_type, cond, new_low, ConstantU32(state, 0)});
state.builder.AddFunction(
{OpSelect, state.uint_type, scc, cond, ConstantU32(state, 1), ConstantU32(state, 0)});
EmitStoreU32(state, SccOperand(), scc);
@@ -269,18 +268,19 @@ void EmitReadFirstLaneU32(EmitterState& state, const IR::Instruction& inst) {
const auto first_lane = state.builder.AllocateId();
const auto first_value = state.builder.AllocateId();
state.builder.AddFunction({OpGroupNonUniformBallot, state.vec4_uint_type, ballot,
ConstantU32(state, ScopeSubgroup), active});
ConstantU32(state, ScopeSubgroup), active});
state.builder.AddFunction({OpGroupNonUniformBallotFindLSB, state.uint_type, first_lane,
ConstantU32(state, ScopeSubgroup), ballot});
ConstantU32(state, ScopeSubgroup), ballot});
state.builder.AddFunction({OpGroupNonUniformShuffle, state.uint_type, first_value,
ConstantU32(state, ScopeSubgroup), src, first_lane});
ConstantU32(state, ScopeSubgroup), src, first_lane});
EmitStoreU32(state, inst.dst, first_value);
}
uint32_t EmitLaneIndex(EmitterState& state, const IR::Operand& operand) {
const auto lane = state.builder.AllocateId();
const auto mask = state.wave_size == 32u ? 31u : 63u;
state.builder.AddFunction({OpBitwiseAnd, state.uint_type, lane, EmitValueLoad(state, operand),
ConstantU32(state, 63)});
ConstantU32(state, mask)});
return lane;
}
@@ -289,7 +289,7 @@ void EmitReadLaneU32(EmitterState& state, const IR::Instruction& inst) {
const auto lane = EmitLaneIndex(state, inst.src[1]);
const auto value = state.builder.AllocateId();
state.builder.AddFunction({OpGroupNonUniformShuffle, state.uint_type, value,
ConstantU32(state, ScopeSubgroup), src, lane});
ConstantU32(state, ScopeSubgroup), src, lane});
EmitStoreU32(state, inst.dst, value);
}
@@ -336,10 +336,8 @@ void EmitPermlaneB32(EmitterState& state, const IR::Instruction& inst, bool x16)
state.builder.AddFunction(
{OpBitwiseXor, state.uint_type, row_value, row, ConstantU32(state, 16)});
}
state.builder.AddFunction(
{OpBitwiseAnd, state.uint_type, lane, subid, ConstantU32(state, 15)});
state.builder.AddFunction(
{OpBitwiseAnd, state.uint_type, lane8, lane, ConstantU32(state, 7)});
state.builder.AddFunction({OpBitwiseAnd, state.uint_type, lane, subid, ConstantU32(state, 15)});
state.builder.AddFunction({OpBitwiseAnd, state.uint_type, lane8, lane, ConstantU32(state, 7)});
state.builder.AddFunction(
{OpShiftLeftLogical, state.uint_type, shift, lane8, ConstantU32(state, 2)});
state.builder.AddFunction(
@@ -350,7 +348,7 @@ void EmitPermlaneB32(EmitterState& state, const IR::Instruction& inst, bool x16)
{OpBitwiseAnd, state.uint_type, index1, index0, ConstantU32(state, 15)});
state.builder.AddFunction({OpBitwiseOr, state.uint_type, target, row_value, index1});
state.builder.AddFunction({OpGroupNonUniformShuffle, state.uint_type, shuffled,
ConstantU32(state, ScopeSubgroup), value, target});
ConstantU32(state, ScopeSubgroup), value, target});
uint32_t ret = shuffled;
if (!inst.dst.op_sel) {
const auto source_active = EmitLaneIndexActiveBool(state, target);
@@ -375,7 +373,7 @@ void EmitBarrier(EmitterState& state, const IR::Instruction& inst) {
(void)inst;
const auto semantics = MemorySemanticsAcquireRelease | MemorySemanticsWorkgroupMemory;
state.builder.AddFunction({OpControlBarrier, ConstantU32(state, ScopeWorkgroup),
ConstantU32(state, ScopeWorkgroup), ConstantU32(state, semantics)});
ConstantU32(state, ScopeWorkgroup), ConstantU32(state, semantics)});
}
} // namespace Libs::Graphics::ShaderRecompiler::Spirv::Emitter
@@ -2,6 +2,66 @@
#include "graphics/shader/recompiler/emitter/spirvEmitterInternal.h"
namespace Libs::Graphics::ShaderRecompiler::Spirv::Emitter {
namespace {
uint32_t EmitCubeAxisF32(EmitterState& state, uint32_t value) {
const auto normalized = state.builder.AllocateId();
state.builder.AddFunction(
{OpFSub, state.float_type, normalized, value, ConstantF32(state, 0x3f800000u)});
return normalized;
}
uint32_t EmitCubeLayerF32(EmitterState& state, uint32_t face_id) {
// Sampled RDNA2 cubemaps encode face_id as slice * 8 + face. The native
// 2D-array view stores six contiguous faces per slice, so remove the two
// reserved face IDs from every preceding slice.
const auto guest_layer = state.builder.AllocateId();
const auto slice = state.builder.AllocateId();
const auto padding = state.builder.AllocateId();
const auto host_layer = state.builder.AllocateId();
const auto result = state.builder.AllocateId();
state.builder.AddFunction({OpConvertFToU, state.uint_type, guest_layer, face_id});
state.builder.AddFunction(
{OpShiftRightLogical, state.uint_type, slice, guest_layer, ConstantU32(state, 3)});
state.builder.AddFunction(
{OpShiftLeftLogical, state.uint_type, padding, slice, ConstantU32(state, 1)});
state.builder.AddFunction({OpISub, state.uint_type, host_layer, guest_layer, padding});
state.builder.AddFunction({OpConvertUToF, state.float_type, result, host_layer});
return result;
}
uint32_t EmitImageCoordF32Impl(EmitterState& state, const IR::Instruction& inst,
const IR::Operand& address, uint32_t first_component,
uint32_t components) {
auto x = EmitImageAddressFloatLoad(state, inst, address, first_component);
if (components == 1u) {
return x;
}
auto y = inst.memory.image_address_components > first_component + 1u
? EmitImageAddressFloatLoad(state, inst, address, first_component + 1u)
: EmitZeroF32(state);
if (inst.memory.image_cube) {
// RDNA2 sampled cubemap S/T coordinates are biased by +1 relative to
// normalized 2D-array coordinates.
x = EmitCubeAxisF32(state, x);
y = EmitCubeAxisF32(state, y);
}
const auto coord = state.builder.AllocateId();
if (components == 3u) {
auto z = inst.memory.image_address_components > first_component + 2u
? EmitImageAddressFloatLoad(state, inst, address, first_component + 2u)
: EmitZeroF32(state);
if (inst.memory.image_cube) {
z = EmitCubeLayerF32(state, z);
}
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, coord, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, coord, x, y});
}
return coord;
}
} // namespace
bool HasImageSampleFlag(const IR::Instruction& inst, uint32_t flag) {
return (inst.memory.image_sample_flags & flag) != 0;
@@ -21,7 +81,7 @@ ImageSampleLayout MakeImageSampleLayout(const IR::Instruction& inst, ImageViewKi
}
if (HasImageSampleFlag(inst, Decoder::ImageSampleFlagDerivative)) {
const auto components = ImageViewSpatialComponents(view);
layout.grad_x = cursor;
layout.grad_x = cursor;
cursor += components;
layout.grad_y = cursor;
cursor += components;
@@ -36,24 +96,8 @@ ImageSampleLayout MakeImageSampleLayout(const IR::Instruction& inst, ImageViewKi
uint32_t EmitImageCoordF32(EmitterState& state, const IR::Instruction& inst,
const ImageSampleLayout& layout, ImageViewKind view) {
const auto x = EmitImageAddressFloatLoad(state, inst, inst.src[0], layout.coord);
const auto components = ImageViewCoordinateComponents(view);
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > layout.coord + 1u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], layout.coord + 1u)
: EmitZeroF32(state);
const auto coord = state.builder.AllocateId();
if (components == 3u) {
const auto z = inst.memory.image_address_components > layout.coord + 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], layout.coord + 2u)
: EmitZeroF32(state);
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, coord, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, coord, x, y});
}
return coord;
return EmitImageCoordF32Impl(state, inst, inst.src[0], layout.coord,
ImageViewCoordinateComponents(view));
}
uint32_t EmitImageLodF32(EmitterState& state, const IR::Instruction& inst,
@@ -95,10 +139,10 @@ uint32_t EmitImageGradientF32(EmitterState& state, const IR::Instruction& inst,
: EmitZeroF32(state);
const auto grad = state.builder.AllocateId();
if (components == 3u) {
const auto z = inst.memory.image_address_components > first_component + 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0],
first_component + 2u)
: EmitZeroF32(state);
const auto z =
inst.memory.image_address_components > first_component + 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], first_component + 2u)
: EmitZeroF32(state);
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, grad, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, grad, x, y});
@@ -120,8 +164,7 @@ uint32_t EmitImagePackedOffsetI32(EmitterState& state, const IR::Instruction& in
state.builder.AddFunction(
{OpCompositeConstruct, state.vec3_int_type, ret, zero, zero, zero});
} else {
state.builder.AddFunction(
{OpCompositeConstruct, state.vec2_int_type, ret, zero, zero});
state.builder.AddFunction({OpCompositeConstruct, state.vec2_int_type, ret, zero, zero});
}
return ret;
}
@@ -131,18 +174,18 @@ uint32_t EmitImagePackedOffsetI32(EmitterState& state, const IR::Instruction& in
const auto offset_x = state.builder.AllocateId();
state.builder.AddFunction({OpBitcast, state.int_type, packed_i32, packed_bits});
state.builder.AddFunction({OpBitFieldSExtract, state.int_type, offset_x, packed_i32,
ConstantI32(state, 0), ConstantI32(state, 6)});
ConstantI32(state, 0), ConstantI32(state, 6)});
if (components == 1u) {
return offset_x;
}
const auto offset_y = state.builder.AllocateId();
const auto offset = state.builder.AllocateId();
state.builder.AddFunction({OpBitFieldSExtract, state.int_type, offset_y, packed_i32,
ConstantI32(state, 8), ConstantI32(state, 6)});
ConstantI32(state, 8), ConstantI32(state, 6)});
if (components == 3u) {
const auto offset_z = state.builder.AllocateId();
state.builder.AddFunction({OpBitFieldSExtract, state.int_type, offset_z, packed_i32,
ConstantI32(state, 16), ConstantI32(state, 6)});
ConstantI32(state, 16), ConstantI32(state, 6)});
state.builder.AddFunction(
{OpCompositeConstruct, state.vec3_int_type, offset, offset_x, offset_y, offset_z});
} else {
@@ -158,7 +201,7 @@ uint32_t EmitImageCoordU32(EmitterState& state, const IR::Instruction& inst, Ima
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > 1u
const auto y = inst.memory.image_address_components > 1u
? EmitImageAddressValueLoad(state, inst, inst.src[1], 1)
: ConstantU32(state, 0);
const auto coord = state.builder.AllocateId();
@@ -180,7 +223,7 @@ uint32_t EmitImageLoadCoordU32(EmitterState& state, const IR::Instruction& inst,
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > 1u
const auto y = inst.memory.image_address_components > 1u
? EmitImageAddressValueLoad(state, inst, inst.src[0], 1)
: ConstantU32(state, 0);
const auto coord = state.builder.AllocateId();
@@ -209,24 +252,8 @@ uint32_t EmitImageMipLodU32(EmitterState& state, const IR::Instruction& inst,
uint32_t EmitImageQueryCoordF32(EmitterState& state, const IR::Instruction& inst,
ImageViewKind view) {
const auto x = EmitImageAddressFloatLoad(state, inst, inst.src[0], 0);
const auto components = ImageViewCoordinateComponents(view);
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > 1u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], 1)
: EmitZeroF32(state);
const auto coord = state.builder.AllocateId();
if (components == 3u) {
const auto z = inst.memory.image_address_components > 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], 2)
: EmitZeroF32(state);
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, coord, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, coord, x, y});
}
return coord;
// OpImageQueryLod takes only the spatial coordinates, even for arrayed images.
return EmitImageCoordF32Impl(state, inst, inst.src[0], 0, ImageViewSpatialComponents(view));
}
uint32_t DmaskComponentIndex(uint32_t dmask, uint32_t component) {
@@ -33,13 +33,13 @@ uint32_t ConstantImageGatherHorizontalOffsets(EmitterState& state, ImageViewKind
}
uint32_t LoadStorageImageDescriptorAtIndex(EmitterState& state, uint32_t resource,
uint32_t array_index, bool uint_image,
uint32_t array_index_id, bool uint_image,
ImageViewKind view) {
const auto kind = StorageBindingKind(uint_image, view);
const auto kind = StorageBindingKind(uint_image, view);
const auto& descriptors = state.storage_images[StorageImageIndex(uint_image, view)];
const auto pointer =
DescriptorElementPointer(state, descriptors.pointer_type, descriptors.variable, array_index,
kind, resource, "storage image descriptor array was not emitted");
const auto pointer = DescriptorElementPointerId(
state, descriptors.pointer_type, descriptors.variable, array_index_id, kind, resource,
"storage image descriptor array was not emitted");
const auto image = state.builder.AllocateId();
state.builder.AddFunction({OpLoad, descriptors.image_type, image, pointer});
return image;
@@ -133,10 +133,18 @@ void EmitImageLoad(EmitterState& state, const IR::Instruction& inst) {
const bool integer = inst.memory.kind == IR::ResourceKind::ImageUint;
const auto color = state.builder.AllocateId();
state.builder.AddFunction({OpImageFetch, integer ? state.vec4_uint_type : state.vec4_float_type,
color, image, EmitImageLoadCoordU32(state, inst, view),
ImageOperandsLodMask,
EmitImageMipLodU32(state, inst, inst.src[0], view)});
const auto coord = EmitImageLoadCoordU32(state, inst, view);
if (ImageSpirvMultisampled(view) != 0) {
const auto sample = EmitImageAddressValueLoad(state, inst, inst.src[0],
ImageViewCoordinateComponents(view));
state.builder.AddFunction({OpImageFetch,
integer ? state.vec4_uint_type : state.vec4_float_type, color,
image, coord, ImageOperandsSampleMask, sample});
} else {
state.builder.AddFunction(
{OpImageFetch, integer ? state.vec4_uint_type : state.vec4_float_type, color, image,
coord, ImageOperandsLodMask, EmitImageMipLodU32(state, inst, inst.src[0], view)});
}
const auto dmask = inst.memory.dmask != 0 ? inst.memory.dmask : 1u;
uint32_t dst_index = 0;
@@ -158,14 +166,40 @@ void EmitImageLoad(EmitterState& state, const IR::Instruction& inst) {
void EmitImageStore(EmitterState& state, const IR::Instruction& inst) {
const auto uint_image = inst.memory.kind == IR::ResourceKind::StorageImageUint;
const auto view = StorageImageViewKind(state, inst.memory, uint_image, inst.pc);
const auto binding = ResourceForDescriptor(state, StorageBindingKind(uint_image, view),
inst.memory.resource);
const auto image = LoadStorageImageDescriptorAtIndex(state, inst.memory.resource,
binding.array_index, uint_image, view);
const auto binding =
ResourceForDescriptor(state, StorageBindingKind(uint_image, view), inst.memory.resource);
const auto emit_write = [&](uint32_t descriptor_index, bool non_uniform) {
if (non_uniform) {
state.builder.AddAnnotation({OpDecorate, descriptor_index, DecorationNonUniform});
}
const auto image = LoadStorageImageDescriptorAtIndex(state, inst.memory.resource,
descriptor_index, uint_image, view);
if (non_uniform) {
state.builder.AddAnnotation({OpDecorate, image, DecorationNonUniform});
}
state.builder.AddFunction({OpImageWrite, image, EmitImageCoordU32(state, inst, view),
uint_image ? EmitImageStoreTexelU32(state, inst)
: EmitImageStoreTexelF32(state, inst)});
};
if (!inst.memory.image_has_mip) {
emit_write(ConstantU32(state, binding.array_index), false);
return;
}
const auto& resource = state.program.info.images[inst.memory.resource];
const auto mip = EmitImageMipLodU32(state, inst, inst.src[1], view);
const auto in_range = state.builder.AllocateId();
state.builder.AddFunction(
{OpImageWrite, image, EmitImageCoordU32(state, inst, view),
uint_image ? EmitImageStoreTexelU32(state, inst) : EmitImageStoreTexelF32(state, inst)});
{OpULessThan, state.bool_type, in_range, mip, ConstantU32(state, resource.mip_levels)});
EmitIfCondition(state, in_range, [&] {
auto descriptor_index = mip;
if (binding.array_index != 0) {
descriptor_index = state.builder.AllocateId();
state.builder.AddFunction({OpIAdd, state.uint_type, descriptor_index,
ConstantU32(state, binding.array_index), mip});
}
emit_write(descriptor_index, true);
});
}
void EmitImageSampleResult(EmitterState& state, const IR::Instruction& inst, uint32_t sample,
@@ -261,9 +295,9 @@ void EmitImageSample(EmitterState& state, const IR::Instruction& inst) {
} else if (integer) {
result_type = state.vec4_uint_type;
}
const auto explicit_lod = ImageSampleNeedsExplicitLod(state, inst);
const auto opcode = ImageSampleOpcode(state, inst);
std::vector<uint32_t> words = {opcode, result_type, sample, sampled_image, base_coord};
const auto explicit_lod = ImageSampleNeedsExplicitLod(state, inst);
const auto opcode = ImageSampleOpcode(state, inst);
std::vector<uint32_t> words = {opcode, result_type, sample, sampled_image, base_coord};
if (dref) {
words.push_back(EmitImageDrefF32(state, inst, layout));
}
@@ -23,38 +23,40 @@
namespace Libs::Graphics::ShaderRecompiler::Spirv::Emitter {
enum : uint32_t {
ExecutionModelVertex = 0,
ExecutionModelFragment = 4,
ExecutionModelGLCompute = 5,
ExecutionModeOriginUpperLeft = 7,
ExecutionModeEarlyFragmentTests = 9,
ExecutionModeDepthReplacing = 12,
ExecutionModeLocalSize = 17,
ExecutionModeDerivativeGroupQuadsKHR = 5289,
AddressingModelLogical = 0,
MemoryModelGLSL450 = 1,
CapabilityShader = 1,
CapabilityImageGatherExtended = 25,
CapabilitySampled1D = 43,
CapabilityImage1D = 44,
CapabilityImageQuery = 50,
CapabilityStorageImageReadWithoutFormat = 55,
CapabilityStorageImageWriteWithoutFormat = 56,
CapabilityGroupNonUniform = 61,
CapabilityGroupNonUniformBallot = 64,
CapabilityGroupNonUniformShuffle = 65,
CapabilityComputeDerivativeGroupQuadsKHR = 5288,
StorageClassUniformConstant = 0,
StorageClassInput = 1,
StorageClassOutput = 3,
StorageClassWorkgroup = 4,
StorageClassFunction = 7,
StorageClassPushConstant = 9,
StorageClassImage = 11,
StorageClassStorageBuffer = 12,
FunctionControlNone = 0,
SelectionControlNone = 0,
LoopControlNone = 0,
ExecutionModelVertex = 0,
ExecutionModelFragment = 4,
ExecutionModelGLCompute = 5,
ExecutionModeOriginUpperLeft = 7,
ExecutionModeEarlyFragmentTests = 9,
ExecutionModeDepthReplacing = 12,
ExecutionModeLocalSize = 17,
ExecutionModeDerivativeGroupQuadsKHR = 5289,
AddressingModelLogical = 0,
MemoryModelGLSL450 = 1,
CapabilityShader = 1,
CapabilityImageGatherExtended = 25,
CapabilitySampled1D = 43,
CapabilityImage1D = 44,
CapabilityImageQuery = 50,
CapabilityStorageImageReadWithoutFormat = 55,
CapabilityStorageImageWriteWithoutFormat = 56,
CapabilityGroupNonUniform = 61,
CapabilityGroupNonUniformBallot = 64,
CapabilityGroupNonUniformShuffle = 65,
CapabilityShaderNonUniform = 5301,
CapabilityStorageImageArrayNonUniformIndexing = 5309,
CapabilityComputeDerivativeGroupQuadsKHR = 5288,
StorageClassUniformConstant = 0,
StorageClassInput = 1,
StorageClassOutput = 3,
StorageClassWorkgroup = 4,
StorageClassFunction = 7,
StorageClassPushConstant = 9,
StorageClassImage = 11,
StorageClassStorageBuffer = 12,
FunctionControlNone = 0,
SelectionControlNone = 0,
LoopControlNone = 0,
};
enum : uint32_t {
@@ -67,6 +69,7 @@ enum : uint32_t {
DecorationBinding = 33,
DecorationDescriptorSet = 34,
DecorationOffset = 35,
DecorationNonUniform = 5300,
};
enum : uint32_t {
@@ -99,6 +102,7 @@ enum : uint32_t {
ImageOperandsGradMask = 0x00000004u,
ImageOperandsOffsetMask = 0x00000010u,
ImageOperandsConstOffsetsMask = 0x00000020u,
ImageOperandsSampleMask = 0x00000040u,
};
enum : uint32_t {
@@ -150,7 +154,6 @@ enum : uint32_t {
OpImageGather = 96,
OpImageDrefGather = 97,
OpImageWrite = 99,
OpImage = 100,
OpImageQuerySizeLod = 103,
OpImageQueryLod = 105,
OpImageQueryLevels = 106,
@@ -382,7 +385,7 @@ struct EmitterState {
uint32_t ptr_workgroup_array = 0;
uint32_t ptr_workgroup_uint = 0;
uint32_t lds_variable = 0;
std::array<SampledImageDescriptors, 10> sampled_images;
std::array<SampledImageDescriptors, 14> sampled_images;
std::array<StorageImageDescriptors, 10> storage_images;
uint32_t sampler_type = 0;
uint32_t sampler_array_type = 0;
@@ -453,17 +456,20 @@ enum class ImageViewKind {
Dim2D,
Dim2DArray,
Dim3D,
Dim2DMsaa,
Dim2DMsaaArray,
Count,
};
constexpr uint32_t ImageViewKindCount = static_cast<uint32_t>(ImageViewKind::Count);
constexpr uint32_t SampledImageViewKindCount = static_cast<uint32_t>(ImageViewKind::Count);
constexpr uint32_t StorageImageViewKindCount = static_cast<uint32_t>(ImageViewKind::Dim2DMsaa);
constexpr uint32_t SampledImageIndex(bool integer, ImageViewKind view) {
return static_cast<uint32_t>(view) + (integer ? ImageViewKindCount : 0u);
return static_cast<uint32_t>(view) + (integer ? SampledImageViewKindCount : 0u);
}
constexpr uint32_t StorageImageIndex(bool integer, ImageViewKind view) {
return static_cast<uint32_t>(view) + (integer ? ImageViewKindCount : 0u);
return static_cast<uint32_t>(view) + (integer ? StorageImageViewKindCount : 0u);
}
constexpr IR::DescriptorBindingKind SampledBindingKind(bool integer, ImageViewKind view) {
@@ -474,6 +480,9 @@ constexpr IR::DescriptorBindingKind SampledBindingKind(bool integer, ImageViewKi
case ImageViewKind::Dim2D: return IR::DescriptorBindingKind::SampledUint2D;
case ImageViewKind::Dim2DArray: return IR::DescriptorBindingKind::SampledUint2DArray;
case ImageViewKind::Dim3D: return IR::DescriptorBindingKind::SampledUint3D;
case ImageViewKind::Dim2DMsaa: return IR::DescriptorBindingKind::SampledUint2DMsaa;
case ImageViewKind::Dim2DMsaaArray:
return IR::DescriptorBindingKind::SampledUint2DMsaaArray;
default: break;
}
}
@@ -483,6 +492,8 @@ constexpr IR::DescriptorBindingKind SampledBindingKind(bool integer, ImageViewKi
case ImageViewKind::Dim2D: return IR::DescriptorBindingKind::Sampled2D;
case ImageViewKind::Dim2DArray: return IR::DescriptorBindingKind::Sampled2DArray;
case ImageViewKind::Dim3D: return IR::DescriptorBindingKind::Sampled3D;
case ImageViewKind::Dim2DMsaa: return IR::DescriptorBindingKind::Sampled2DMsaa;
case ImageViewKind::Dim2DMsaaArray: return IR::DescriptorBindingKind::Sampled2DMsaaArray;
default: break;
}
return IR::DescriptorBindingKind::Count;
@@ -516,6 +527,8 @@ constexpr uint32_t ImageSpirvDimension(ImageViewKind view) {
case ImageViewKind::Dim1DArray: return Dim1D;
case ImageViewKind::Dim2D:
case ImageViewKind::Dim2DArray:
case ImageViewKind::Dim2DMsaa:
case ImageViewKind::Dim2DMsaaArray:
case ImageViewKind::Count: return Dim2D;
case ImageViewKind::Dim3D: return Dim3D;
}
@@ -523,7 +536,14 @@ constexpr uint32_t ImageSpirvDimension(ImageViewKind view) {
}
constexpr uint32_t ImageSpirvArrayed(ImageViewKind view) {
return view == ImageViewKind::Dim1DArray || view == ImageViewKind::Dim2DArray ? 1u : 0u;
return view == ImageViewKind::Dim1DArray || view == ImageViewKind::Dim2DArray ||
view == ImageViewKind::Dim2DMsaaArray
? 1u
: 0u;
}
constexpr uint32_t ImageSpirvMultisampled(ImageViewKind view) {
return view == ImageViewKind::Dim2DMsaa || view == ImageViewKind::Dim2DMsaaArray ? 1u : 0u;
}
struct AddCarryResult {
@@ -662,6 +682,10 @@ uint32_t DescriptorElementPointer(EmitterState& state, uint32_t result_ptr_type,
uint32_t variable_id, uint32_t array_index,
IR::DescriptorBindingKind kind, uint32_t resource,
const char* variable_name);
uint32_t DescriptorElementPointerId(EmitterState& state, uint32_t result_ptr_type,
uint32_t variable_id, uint32_t array_index_id,
IR::DescriptorBindingKind kind, uint32_t resource,
const char* variable_name);
ImageViewKind SampledImageViewKind(const EmitterState& state, const IR::MemoryInfo& mem,
uint32_t use_pc);
@@ -174,6 +174,12 @@ uint32_t VertexParameterInputPointerType(const EmitterState& state, VertexInputS
}
}
static bool MrtUsesUintOutput(const EmitterState& state, uint32_t index) {
return state.stage == ShaderType::Pixel && state.pixel_input_info != nullptr &&
index < std::size(state.pixel_input_info->target_output_mode) &&
state.pixel_input_info->target_output_mode[index] == 7u;
}
void AllocateInputVariables(EmitterState& state) {
for (auto& binding: state.inputs) {
binding.variable_id = state.builder.AllocateId();
@@ -323,23 +329,39 @@ void AddDescriptorAnnotationsAndNames(EmitterState& state) {
Decorate(state.address_memory_variable, "address_memory",
IR::DescriptorBindingKind::AddressMemory);
}
constexpr const char* SampledNames[] = {
"sampled_1d", "sampled_1d_array", "sampled_2d", "sampled_2d_array",
"sampled_3d", "sampled_uint_1d", "sampled_uint_1d_array",
"sampled_uint_2d", "sampled_uint_2d_array", "sampled_uint_3d"};
constexpr const char* SampledNames[] = {"sampled_1d",
"sampled_1d_array",
"sampled_2d",
"sampled_2d_array",
"sampled_3d",
"sampled_2d_msaa",
"sampled_2d_msaa_array",
"sampled_uint_1d",
"sampled_uint_1d_array",
"sampled_uint_2d",
"sampled_uint_2d_array",
"sampled_uint_3d",
"sampled_uint_2d_msaa",
"sampled_uint_2d_msaa_array"};
for (uint32_t i = 0; i < state.sampled_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const auto view = static_cast<ImageViewKind>(i % SampledImageViewKindCount);
Decorate(state.sampled_images[i].variable, SampledNames[i],
SampledBindingKind(i >= ImageViewKindCount, view));
SampledBindingKind(i >= SampledImageViewKindCount, view));
}
constexpr const char* StorageNames[] = {
"storage_1d", "storage_1d_array", "storage_2d", "storage_2d_array",
"storage_3d", "storage_uint_1d", "storage_uint_1d_array",
"storage_uint_2d", "storage_uint_2d_array", "storage_uint_3d"};
constexpr const char* StorageNames[] = {"storage_1d",
"storage_1d_array",
"storage_2d",
"storage_2d_array",
"storage_3d",
"storage_uint_1d",
"storage_uint_1d_array",
"storage_uint_2d",
"storage_uint_2d_array",
"storage_uint_3d"};
for (uint32_t i = 0; i < state.storage_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const auto view = static_cast<ImageViewKind>(i % StorageImageViewKindCount);
Decorate(state.storage_images[i].variable, StorageNames[i],
StorageBindingKind(i >= ImageViewKindCount, view));
StorageBindingKind(i >= StorageImageViewKindCount, view));
}
if (state.sampler_variable != 0) {
Decorate(state.sampler_variable, "samplers", IR::DescriptorBindingKind::Samplers);
@@ -409,6 +431,7 @@ void EmitHeaderAndTypes(EmitterState& state) {
state.ptr_output_sample_mask_array = state.builder.AllocateId();
state.ptr_output_float = state.builder.AllocateId();
state.ptr_output_vec4_float = state.builder.AllocateId();
const auto ptr_output_vec4_uint = state.builder.AllocateId();
state.per_vertex_type = state.builder.AllocateId();
state.ptr_output_per_vertex = state.builder.AllocateId();
state.storage_runtime_array_type = state.builder.AllocateId();
@@ -444,15 +467,15 @@ void EmitHeaderAndTypes(EmitterState& state) {
image.array_type = state.builder.AllocateId();
image.array_pointer_type = state.builder.AllocateId();
}
state.sampler_type = state.builder.AllocateId();
state.sampler_array_type = state.builder.AllocateId();
state.ptr_uniform_sampler = state.builder.AllocateId();
state.ptr_uniform_sampler_array = state.builder.AllocateId();
state.ptr_image_uint = state.builder.AllocateId();
state.func_type = state.builder.AllocateId();
state.main_func = state.builder.AllocateId();
state.entry_label = state.builder.AllocateId();
state.glsl_std450 = state.builder.AllocateId();
state.sampler_type = state.builder.AllocateId();
state.sampler_array_type = state.builder.AllocateId();
state.ptr_uniform_sampler = state.builder.AllocateId();
state.ptr_uniform_sampler_array = state.builder.AllocateId();
state.ptr_image_uint = state.builder.AllocateId();
state.func_type = state.builder.AllocateId();
state.main_func = state.builder.AllocateId();
state.entry_label = state.builder.AllocateId();
state.glsl_std450 = state.builder.AllocateId();
state.builder.AddCapability({CapabilityShader});
state.builder.AddCapability({CapabilitySampled1D});
@@ -461,8 +484,15 @@ void EmitHeaderAndTypes(EmitterState& state) {
if (state.needs_image_gather_extended) {
state.builder.AddCapability({CapabilityImageGatherExtended});
}
if (std::any_of(
state.program.info.images.begin(), state.program.info.images.end(),
[](const auto& image) { return image.mip_mode == IR::ImageMipMode::DynamicStorage; })) {
state.builder.AddCapability({CapabilityShaderNonUniform});
state.builder.AddCapability({CapabilityStorageImageArrayNonUniformIndexing});
state.builder.AddExtension("SPV_EXT_descriptor_indexing");
}
if (std::any_of(state.storage_images.begin(),
state.storage_images.begin() + ImageViewKindCount,
state.storage_images.begin() + StorageImageViewKindCount,
[](const auto& image) { return image.variable != 0; })) {
state.builder.AddCapability({CapabilityStorageImageReadWithoutFormat});
state.builder.AddCapability({CapabilityStorageImageWriteWithoutFormat});
@@ -605,6 +635,8 @@ void EmitHeaderAndTypes(EmitterState& state) {
{OpTypePointer, state.ptr_output_int, StorageClassOutput, state.int_type});
state.builder.AddType(
{OpTypePointer, state.ptr_output_vec4_float, StorageClassOutput, state.vec4_float_type});
state.builder.AddType(
{OpTypePointer, ptr_output_vec4_uint, StorageClassOutput, state.vec4_uint_type});
if (state.per_vertex_variable != 0) {
state.builder.AddType({OpTypeStruct, state.per_vertex_type, state.vec4_float_type});
state.builder.AddType({OpTypePointer, state.ptr_output_per_vertex, StorageClassOutput,
@@ -615,8 +647,12 @@ void EmitHeaderAndTypes(EmitterState& state) {
for (const auto& binding: state.outputs) {
if (binding.kind == IR::StageOutputKind::Parameter ||
binding.kind == IR::StageOutputKind::Mrt) {
const auto pointer_type =
binding.kind == IR::StageOutputKind::Mrt && MrtUsesUintOutput(state, binding.index)
? ptr_output_vec4_uint
: state.ptr_output_vec4_float;
state.builder.AddType(
{OpVariable, state.ptr_output_vec4_float, binding.variable_id, StorageClassOutput});
{OpVariable, pointer_type, binding.variable_id, StorageClassOutput});
}
}
if (state.depth_variable != 0) {
@@ -700,11 +736,11 @@ void EmitHeaderAndTypes(EmitterState& state) {
}
for (uint32_t i = 0; i < state.sampled_images.size(); i++) {
auto& image = state.sampled_images[i];
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const bool integer = i >= ImageViewKindCount;
const auto view = static_cast<ImageViewKind>(i % SampledImageViewKindCount);
const bool integer = i >= SampledImageViewKindCount;
const auto component = integer ? state.uint_type : state.float_type;
state.builder.AddType({OpTypeImage, image.image_type, component,
ImageSpirvDimension(view), 0, ImageSpirvArrayed(view), 0, 1,
state.builder.AddType({OpTypeImage, image.image_type, component, ImageSpirvDimension(view),
0, ImageSpirvArrayed(view), ImageSpirvMultisampled(view), 1,
ImageFormatUnknown});
state.builder.AddType({OpTypeSampledImage, image.sampled_image_type, image.image_type});
state.builder.AddType(
@@ -733,13 +769,12 @@ void EmitHeaderAndTypes(EmitterState& state) {
}
for (uint32_t i = 0; i < state.storage_images.size(); i++) {
auto& image = state.storage_images[i];
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const bool integer = i >= ImageViewKindCount;
const auto view = static_cast<ImageViewKind>(i % StorageImageViewKindCount);
const bool integer = i >= StorageImageViewKindCount;
const auto component = integer ? state.uint_type : state.float_type;
const auto format = integer ? ImageFormatR32ui : ImageFormatUnknown;
state.builder.AddType({OpTypeImage, image.image_type, component,
ImageSpirvDimension(view), 0, ImageSpirvArrayed(view), 0, 2,
format});
state.builder.AddType({OpTypeImage, image.image_type, component, ImageSpirvDimension(view),
0, ImageSpirvArrayed(view), 0, 2, format});
state.builder.AddType(
{OpTypePointer, image.pointer_type, StorageClassUniformConstant, image.image_type});
if (image.variable != 0) {
@@ -786,15 +821,15 @@ void AllocateDescriptorVariables(EmitterState& state) {
state.flattened_srt_variable = state.builder.AllocateId();
}
for (uint32_t i = 0; i < state.sampled_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
if (DescriptorBinding(state, SampledBindingKind(i >= ImageViewKindCount, view)) !=
const auto view = static_cast<ImageViewKind>(i % SampledImageViewKindCount);
if (DescriptorBinding(state, SampledBindingKind(i >= SampledImageViewKindCount, view)) !=
nullptr) {
state.sampled_images[i].variable = state.builder.AllocateId();
}
}
for (uint32_t i = 0; i < state.storage_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
if (DescriptorBinding(state, StorageBindingKind(i >= ImageViewKindCount, view)) !=
const auto view = static_cast<ImageViewKind>(i % StorageImageViewKindCount);
if (DescriptorBinding(state, StorageBindingKind(i >= StorageImageViewKindCount, view)) !=
nullptr) {
state.storage_images[i].variable = state.builder.AllocateId();
}
@@ -13,16 +13,30 @@ namespace {
constexpr uint32_t MaxPushConstantBytes = 128;
constexpr std::array ImageBindingKinds = {
DescriptorBindingKind::Sampled1D, DescriptorBindingKind::Sampled1DArray,
DescriptorBindingKind::Sampled2D, DescriptorBindingKind::Sampled2DArray,
DescriptorBindingKind::Sampled3D, DescriptorBindingKind::SampledUint1D,
DescriptorBindingKind::SampledUint1DArray, DescriptorBindingKind::SampledUint2D,
DescriptorBindingKind::SampledUint2DArray, DescriptorBindingKind::SampledUint3D,
DescriptorBindingKind::Storage1D, DescriptorBindingKind::Storage1DArray,
DescriptorBindingKind::Storage2D, DescriptorBindingKind::Storage2DArray,
DescriptorBindingKind::Storage3D, DescriptorBindingKind::StorageUint1D,
DescriptorBindingKind::StorageUint1DArray, DescriptorBindingKind::StorageUint2D,
DescriptorBindingKind::StorageUint2DArray, DescriptorBindingKind::StorageUint3D,
DescriptorBindingKind::Sampled1D,
DescriptorBindingKind::Sampled1DArray,
DescriptorBindingKind::Sampled2D,
DescriptorBindingKind::Sampled2DArray,
DescriptorBindingKind::Sampled2DMsaa,
DescriptorBindingKind::Sampled2DMsaaArray,
DescriptorBindingKind::Sampled3D,
DescriptorBindingKind::SampledUint1D,
DescriptorBindingKind::SampledUint1DArray,
DescriptorBindingKind::SampledUint2D,
DescriptorBindingKind::SampledUint2DArray,
DescriptorBindingKind::SampledUint2DMsaa,
DescriptorBindingKind::SampledUint2DMsaaArray,
DescriptorBindingKind::SampledUint3D,
DescriptorBindingKind::Storage1D,
DescriptorBindingKind::Storage1DArray,
DescriptorBindingKind::Storage2D,
DescriptorBindingKind::Storage2DArray,
DescriptorBindingKind::Storage3D,
DescriptorBindingKind::StorageUint1D,
DescriptorBindingKind::StorageUint1DArray,
DescriptorBindingKind::StorageUint2D,
DescriptorBindingKind::StorageUint2DArray,
DescriptorBindingKind::StorageUint3D,
};
bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
@@ -36,6 +50,8 @@ bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
case Dimension::Dim1DArray: result = Kind::Sampled1DArray; return true;
case Dimension::Dim2D: result = Kind::Sampled2D; return true;
case Dimension::Dim2DArray: result = Kind::Sampled2DArray; return true;
case Dimension::Dim2DMsaa: result = Kind::Sampled2DMsaa; return true;
case Dimension::Dim2DMsaaArray: result = Kind::Sampled2DMsaaArray; return true;
case Dimension::Dim3D: result = Kind::Sampled3D; return true;
default: return false;
}
@@ -45,6 +61,8 @@ bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
case Dimension::Dim1DArray: result = Kind::SampledUint1DArray; return true;
case Dimension::Dim2D: result = Kind::SampledUint2D; return true;
case Dimension::Dim2DArray: result = Kind::SampledUint2DArray; return true;
case Dimension::Dim2DMsaa: result = Kind::SampledUint2DMsaa; return true;
case Dimension::Dim2DMsaaArray: result = Kind::SampledUint2DMsaaArray; return true;
case Dimension::Dim3D: result = Kind::SampledUint3D; return true;
default: return false;
}
@@ -71,7 +89,7 @@ bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
}
bool CollectValue(const ScalarProvenance& provenance, uint32_t id, std::vector<uint8_t>& visited,
std::set<uint32_t>& registers) {
std::set<uint32_t>& registers) {
if (id <= ScalarProvenance::Unknown) {
return true;
}
@@ -104,7 +122,7 @@ bool CollectValue(const ScalarProvenance& provenance, uint32_t id, std::vector<u
}
bool CollectSource(const Program& program, uint32_t source, bool allow_unknown,
std::vector<uint8_t>& visited, std::set<uint32_t>& registers) {
std::vector<uint8_t>& visited, std::set<uint32_t>& registers) {
if (allow_unknown && source == ScalarProvenance::Unknown) {
return true;
}
@@ -165,8 +183,7 @@ bool CollectUserData(const Program& program, std::vector<uint32_t>& result) {
return false;
}
for (uint32_t i = 0; i < inst.src_count; i++) {
if (!CollectValue(program.provenance, inst.scalar_sources[i], visited,
registers)) {
if (!CollectValue(program.provenance, inst.scalar_sources[i], visited, registers)) {
return false;
}
}
@@ -199,7 +216,7 @@ bool AllocateBindings(Program& program, const BindingLayoutOptions& options, std
if (!program.shader_info_complete || program.binding_layout_complete) {
if (error != nullptr) {
*error = !program.shader_info_complete ? "shader info is not ready"
: "binding layout already allocated";
: "binding layout already allocated";
}
return false;
}
@@ -252,7 +269,10 @@ bool AllocateBindings(Program& program, const BindingLayoutOptions& options, std
}
return false;
}
image_groups[static_cast<size_t>(group - ImageBindingKinds.begin())].push_back(i);
auto& resources = image_groups[static_cast<size_t>(group - ImageBindingKinds.begin())];
for (uint32_t binding = 0; binding < program.info.images[i].NumBindings(); binding++) {
resources.push_back(i);
}
}
for (uint32_t i = 0; i < image_groups.size(); i++) {
if (!image_groups[i].empty()) {
@@ -0,0 +1,323 @@
#include "graphics/shader/recompiler/ir/ReadLaneElimination.h"
#include "graphics/shader/recompiler/ir/SrtWalker.h"
#include <algorithm>
#include <iterator>
#include <map>
#include <set>
#include <utility>
namespace Libs::Graphics::ShaderRecompiler::IR {
namespace {
constexpr uint32_t FirstTemporaryScalarRegister = 128;
struct LaneKey {
uint32_t reg = 0;
uint32_t lane = 0;
auto operator<=>(const LaneKey&) const = default;
};
using LaneSet = std::set<LaneKey>;
bool PairDwordOpcode(Opcode op) {
switch (op) {
case Opcode::MoveU64:
case Opcode::WqmB64:
case Opcode::SaveexecB64:
case Opcode::BitwiseAndU64:
case Opcode::BitwiseAndNotU64:
case Opcode::BitwiseOrU64:
case Opcode::BitwiseOrNotU64:
case Opcode::BitwiseXorU64:
case Opcode::BitwiseNandU64:
case Opcode::BitwiseNorU64:
case Opcode::BitwiseXnorU64:
case Opcode::BitwiseNotU64:
case Opcode::BitFieldMaskU64:
case Opcode::BitFieldExtractU64:
case Opcode::BitReplicateB64B32:
case Opcode::ShiftLeftLogicalU64:
case Opcode::ShiftRightLogicalU64:
case Opcode::SelectU64: return true;
default: return false;
}
}
bool ResolveLane(const Program& program, const Instruction& inst, uint32_t source_index,
uint32_t& lane) {
if (source_index >= inst.src_count || (program.wave_size != 32 && program.wave_size != 64)) {
return false;
}
const auto& selector = inst.src[source_index];
if (selector.kind == OperandKind::ImmediateU32) {
lane = selector.imm % program.wave_size;
return true;
}
uint32_t folded = 0;
if (!FoldScalarConstant(program.provenance, inst.scalar_sources[source_index], folded)) {
return false;
}
lane = folded % program.wave_size;
return true;
}
bool UniformWriteSource(const Instruction& inst) {
if (inst.src_count == 0) {
return false;
}
const auto& source = inst.src[0];
if (source.kind == OperandKind::ImmediateU32 || source.kind == OperandKind::PcRelativeU32) {
return true;
}
return source.kind == OperandKind::Register &&
(source.reg.file == RegisterFile::Scalar || source.reg.file == RegisterFile::Scc ||
source.reg.file == RegisterFile::M0);
}
bool WriteLaneKey(const Program& program, const Instruction& inst, LaneKey& key) {
if (inst.op != Opcode::WriteLaneU32 || inst.dst.kind != OperandKind::Register ||
inst.dst.reg.file != RegisterFile::Vector || !UniformWriteSource(inst)) {
return false;
}
uint32_t lane = 0;
if (!ResolveLane(program, inst, 1, lane)) {
return false;
}
key = {inst.dst.reg.index, lane};
return true;
}
bool ReadLaneKey(const Program& program, const Instruction& inst, LaneKey& key) {
if (inst.op != Opcode::ReadLaneU32 || inst.src_count < 2 ||
inst.src[0].kind != OperandKind::Register || inst.src[0].reg.file != RegisterFile::Vector) {
return false;
}
uint32_t lane = 0;
if (!ResolveLane(program, inst, 1, lane)) {
return false;
}
key = {inst.src[0].reg.index, lane};
return true;
}
void InvalidateRegister(LaneSet& valid, uint32_t reg) {
const auto first = valid.lower_bound({reg, 0});
const auto last = valid.lower_bound({reg + 1u, 0});
valid.erase(first, last);
}
void ApplyInstruction(const Program& program, const Instruction& inst, LaneSet& valid) {
if (inst.op == Opcode::WriteLaneU32 && inst.dst.kind == OperandKind::Register &&
inst.dst.reg.file == RegisterFile::Vector) {
LaneKey key;
if (WriteLaneKey(program, inst, key)) {
valid.insert(key);
return;
}
uint32_t lane = 0;
if (ResolveLane(program, inst, 1, lane)) {
valid.erase({inst.dst.reg.index, lane});
} else {
InvalidateRegister(valid, inst.dst.reg.index);
}
return;
}
if (inst.op == Opcode::MoveRelDestU32 && inst.dst.kind == OperandKind::Register &&
inst.dst.reg.file == RegisterFile::Vector) {
valid.clear();
return;
}
if (inst.dst.kind == OperandKind::Register && inst.dst.reg.file == RegisterFile::Vector) {
uint32_t dwords = std::max(inst.memory.data_dwords, 1u);
if (PairDwordOpcode(inst.op) || inst.op == Opcode::UMadU64U32) {
dwords = std::max(dwords, 2u);
}
for (uint32_t i = 0; i < dwords && inst.dst.reg.index <= UINT32_MAX - i; i++) {
InvalidateRegister(valid, inst.dst.reg.index + i);
}
}
if (inst.dst2.kind == OperandKind::Register && inst.dst2.reg.file == RegisterFile::Vector) {
InvalidateRegister(valid, inst.dst2.reg.index);
}
}
LaneSet TransferBlock(const Program& program, const BasicBlock& block, LaneSet state) {
for (const auto& inst: block.instructions) {
ApplyInstruction(program, inst, state);
}
return state;
}
LaneSet Intersect(const LaneSet& left, const LaneSet& right) {
LaneSet result;
std::set_intersection(left.begin(), left.end(), right.begin(), right.end(),
std::inserter(result, result.end()));
return result;
}
uint32_t NextTemporaryScalarRegister(const Program& program) {
uint32_t next = FirstTemporaryScalarRegister;
const auto consider = [&next](const Operand& operand) {
if (operand.kind == OperandKind::Register && operand.reg.file == RegisterFile::Scalar &&
operand.reg.index >= next && operand.reg.index != UINT32_MAX) {
next = operand.reg.index + 1u;
}
};
for (const auto& block: program.blocks) {
for (const auto& inst: block.instructions) {
consider(inst.dst);
consider(inst.dst2);
for (uint32_t i = 0; i < inst.src_count; i++) {
consider(inst.src[i]);
}
}
}
return next;
}
Operand ScalarRegisterOperand(uint32_t reg) {
Operand operand;
operand.kind = OperandKind::Register;
operand.reg.file = RegisterFile::Scalar;
operand.reg.index = reg;
return operand;
}
Instruction ShadowWrite(const Instruction& write, uint32_t temporary) {
Instruction shadow;
shadow.pc = write.pc;
shadow.op = Opcode::MoveU32;
shadow.dst = ScalarRegisterOperand(temporary);
shadow.src[0] = write.src[0];
shadow.src_count = 1;
return shadow;
}
Instruction ShadowRead(const Instruction& read, uint32_t temporary) {
Instruction rewritten;
rewritten.pc = read.pc;
rewritten.op = Opcode::MoveU32;
rewritten.dst = read.dst;
rewritten.src[0] = ScalarRegisterOperand(temporary);
rewritten.src_count = 1;
return rewritten;
}
} // namespace
ReadLaneEliminationStats EliminateReadLane(Program& program) {
ReadLaneEliminationStats stats;
if (program.blocks.empty() || (program.wave_size != 32 && program.wave_size != 64)) {
return stats;
}
LaneSet universe;
for (const auto& block: program.blocks) {
for (const auto& inst: block.instructions) {
LaneKey key;
if (WriteLaneKey(program, inst, key)) {
universe.insert(key);
}
}
}
if (universe.empty()) {
return stats;
}
const size_t block_count = program.blocks.size();
std::vector<LaneSet> entry(block_count, universe);
std::vector<LaneSet> exit(block_count, universe);
entry[0].clear();
for (size_t block = 0; block < block_count; block++) {
exit[block] = TransferBlock(program, program.blocks[block], entry[block]);
}
bool changed = true;
while (changed) {
changed = false;
for (size_t block_index = 0; block_index < block_count; block_index++) {
LaneSet next_entry;
const auto& block = program.blocks[block_index];
if (block_index != 0 && !block.predecessors.empty()) {
next_entry = universe;
for (const auto predecessor: block.predecessors) {
if (predecessor >= block_count) {
next_entry.clear();
break;
}
next_entry = Intersect(next_entry, exit[predecessor]);
}
}
auto next_exit = TransferBlock(program, block, next_entry);
if (next_entry != entry[block_index] || next_exit != exit[block_index]) {
entry[block_index] = std::move(next_entry);
exit[block_index] = std::move(next_exit);
changed = true;
}
}
}
LaneSet forwarded;
for (size_t block_index = 0; block_index < block_count; block_index++) {
auto state = entry[block_index];
for (const auto& inst: program.blocks[block_index].instructions) {
LaneKey key;
if (ReadLaneKey(program, inst, key) && state.contains(key)) {
forwarded.insert(key);
}
ApplyInstruction(program, inst, state);
}
}
if (forwarded.empty()) {
return stats;
}
std::map<LaneKey, uint32_t> temporaries;
auto next_temporary = NextTemporaryScalarRegister(program);
for (const auto& key: forwarded) {
if (next_temporary == UINT32_MAX) {
return {};
}
temporaries.emplace(key, next_temporary++);
}
for (size_t block_index = 0; block_index < block_count; block_index++) {
const auto original = std::move(program.blocks[block_index].instructions);
auto& rewritten = program.blocks[block_index].instructions;
rewritten.clear();
rewritten.reserve(original.size() + temporaries.size());
auto state = entry[block_index];
for (const auto& inst: original) {
LaneKey read_key;
if (ReadLaneKey(program, inst, read_key) && state.contains(read_key)) {
const auto temporary = temporaries.find(read_key);
if (temporary != temporaries.end()) {
rewritten.push_back(ShadowRead(inst, temporary->second));
stats.rewritten_reads++;
ApplyInstruction(program, inst, state);
continue;
}
}
rewritten.push_back(inst);
LaneKey write_key;
if (WriteLaneKey(program, inst, write_key)) {
const auto temporary = temporaries.find(write_key);
if (temporary != temporaries.end()) {
rewritten.push_back(ShadowWrite(inst, temporary->second));
stats.shadow_writes++;
}
}
ApplyInstruction(program, inst, state);
}
}
return stats;
}
} // namespace Libs::Graphics::ShaderRecompiler::IR
@@ -0,0 +1,20 @@
#ifndef EMULATOR_INCLUDE_EMULATOR_GRAPHICS_SHADER_RECOMPILER_READLANEELIMINATION_H_
#define EMULATOR_INCLUDE_EMULATOR_GRAPHICS_SHADER_RECOMPILER_READLANEELIMINATION_H_
#include "graphics/shader/recompiler/ir/ShaderIR.h"
namespace Libs::Graphics::ShaderRecompiler::IR {
struct ReadLaneEliminationStats {
uint32_t rewritten_reads = 0;
uint32_t shadow_writes = 0;
};
// Replaces fixed-lane ReadLane operations that are reached by a matching WriteLane on every
// control-flow path. A synthetic scalar register snapshots the value at WriteLane execution time,
// so the rewrite remains valid when the source SGPR is subsequently overwritten.
[[nodiscard]] ReadLaneEliminationStats EliminateReadLane(Program& program);
} // namespace Libs::Graphics::ShaderRecompiler::IR
#endif /* EMULATOR_INCLUDE_EMULATOR_GRAPHICS_SHADER_RECOMPILER_READLANEELIMINATION_H_ */
@@ -12,10 +12,11 @@ namespace {
constexpr uint64_t AddressMask = 0x0000ffffffffffffull;
Decoder::ImageDimension DescriptorDimension(const DescriptorValue& descriptor,
Decoder::ImageDimension DescriptorDimension(const DescriptorValue& descriptor,
Decoder::ImageDimension requested) {
const bool is_array = requested == Decoder::ImageDimension::Dim1DArray ||
requested == Decoder::ImageDimension::Dim2DArray;
requested == Decoder::ImageDimension::Dim2DArray ||
requested == Decoder::ImageDimension::Dim2DMsaaArray;
switch (static_cast<Prospero::ImageType>((descriptor.dwords[3] >> 28u) & 0xfu)) {
case Prospero::ImageType::kColor1D: return Decoder::ImageDimension::Dim1D;
case Prospero::ImageType::kColor1DArray:
@@ -26,13 +27,17 @@ Decoder::ImageDimension DescriptorDimension(const DescriptorValue& descrip
case Prospero::ImageType::kColor3D: return Decoder::ImageDimension::Dim3D;
case Prospero::ImageType::kCube: return Decoder::ImageDimension::Dim2DArray;
case Prospero::ImageType::kColor2DArray:
case Prospero::ImageType::kColor2DMsaaArray:
if (is_array) {
return Decoder::ImageDimension::Dim2DArray;
}
return Decoder::ImageDimension::Dim2D;
case Prospero::ImageType::kColor2D:
case Prospero::ImageType::kColor2DMsaa: return Decoder::ImageDimension::Dim2D;
case Prospero::ImageType::kColor2DMsaaArray:
if (is_array) {
return Decoder::ImageDimension::Dim2DMsaaArray;
}
return Decoder::ImageDimension::Dim2DMsaa;
case Prospero::ImageType::kColor2D: return Decoder::ImageDimension::Dim2D;
case Prospero::ImageType::kColor2DMsaa: return Decoder::ImageDimension::Dim2DMsaa;
default: return Decoder::ImageDimension::Unknown;
}
}
@@ -42,8 +47,9 @@ bool NullImageDescriptor(const DescriptorValue& descriptor) {
}
bool ValidImageDescriptor(const DescriptorValue& descriptor) {
const auto type = static_cast<Prospero::ImageType>((descriptor.dwords[3] >> 28u) & 0xfu);
if (type < Prospero::ImageType::kColor1D) {
const auto type = static_cast<Prospero::ImageType>((descriptor.dwords[3] >> 28u) & 0xfu);
const auto format = static_cast<Prospero::BufferFormat>((descriptor.dwords[1] >> 20u) & 0x1ffu);
if (type < Prospero::ImageType::kColor1D || format == Prospero::BufferFormat::kInvalid) {
return false;
}
if (type == Prospero::ImageType::kColor2DMsaa ||
@@ -51,8 +57,7 @@ bool ValidImageDescriptor(const DescriptorValue& descriptor) {
const auto base_level = (descriptor.dwords[3] >> 12u) & 0xfu;
const auto fragments = (descriptor.dwords[3] >> 16u) & 0xfu;
const auto max_mip = (descriptor.dwords[5] >> 4u) & 0xfu;
return base_level == 0 && fragments >= 1 && fragments <= 3 &&
max_mip == fragments;
return base_level == 0 && fragments >= 1 && fragments <= 3 && max_mip == fragments;
}
return true;
}
@@ -61,6 +66,22 @@ uint32_t DescriptorImageSwizzle(const DescriptorValue& descriptor) {
return descriptor.dwords[3] & 0xfffu;
}
bool DescriptorIsCube(const DescriptorValue& descriptor) {
return static_cast<Prospero::ImageType>((descriptor.dwords[3] >> 28u) & 0xfu) ==
Prospero::ImageType::kCube;
}
bool DescriptorMipRange(const DescriptorValue& descriptor, uint32_t& count) {
const auto base_level = (descriptor.dwords[3] >> 12u) & 0xfu;
const auto last_level = (descriptor.dwords[3] >> 16u) & 0xfu;
const auto max_mip = (descriptor.dwords[5] >> 4u) & 0xfu;
if (base_level > last_level || last_level > max_mip) {
return false;
}
count = last_level - base_level + 1u;
return true;
}
bool DecodeBufferDescriptor(const DescriptorValue& descriptor, ShaderBufferResource& result) {
if (descriptor.dword_count != std::size(result.fields)) {
return false;
@@ -171,12 +192,13 @@ bool ValidateResourceSpecialization(const Program& program, const ResourceSnapsh
const auto& image = program.info.images[i];
const auto& descriptor = snapshot.images[i];
if (NullImageDescriptor(descriptor)) {
bool canonical_kind = image.kind == ResourceKind::Image ||
image.kind == ResourceKind::StorageImage;
bool canonical_kind =
image.kind == ResourceKind::Image || image.kind == ResourceKind::StorageImage;
if (image.atomic) {
canonical_kind = image.kind == ResourceKind::StorageImageUint;
}
if (image.dimension != Decoder::ImageDimension::Dim2D || !canonical_kind) {
if (image.dimension != Decoder::ImageDimension::Dim2D || image.cube ||
!canonical_kind) {
if (error != nullptr) {
*error = fmt::format(
"image descriptor {} no longer matches canonical null specialization", i);
@@ -185,11 +207,31 @@ bool ValidateResourceSpecialization(const Program& program, const ResourceSnapsh
}
continue;
}
if (image.mip_mode == ImageMipMode::DynamicStorage) {
uint32_t mip_levels = 0;
if (!DescriptorMipRange(descriptor, mip_levels) || mip_levels != image.mip_levels) {
if (error != nullptr) {
*error = fmt::format(
"image descriptor {} no longer matches specialized storage mip count", i);
}
return false;
}
} else if (image.mip_levels != 1u) {
if (error != nullptr) {
*error = fmt::format("image descriptor {} has invalid non-storage mip count", i);
}
return false;
}
const auto dimension = DescriptorDimension(descriptor, image.dimension);
if (dimension == Decoder::ImageDimension::Unknown || dimension != image.dimension) {
if (dimension == Decoder::ImageDimension::Unknown || dimension != image.dimension ||
DescriptorIsCube(descriptor) != image.cube) {
if (error != nullptr) {
*error =
fmt::format("image descriptor {} no longer matches specialized dimension", i);
fmt::format("image descriptor {} no longer matches specialized dimension: "
"{:08x},{:08x},{:08x},{:08x},{:08x},{:08x},{:08x},{:08x}",
i, descriptor.dwords[0], descriptor.dwords[1], descriptor.dwords[2],
descriptor.dwords[3], descriptor.dwords[4], descriptor.dwords[5],
descriptor.dwords[6], descriptor.dwords[7]);
}
return false;
}
@@ -204,10 +246,13 @@ bool ValidateResourceSpecialization(const Program& program, const ResourceSnapsh
}
return false;
}
const auto uint_descriptor =
Prospero::IsUintTextureFormat((descriptor.dwords[1] >> 20u) & 0x1ffu);
const auto uint_program = image.kind == ResourceKind::ImageUint ||
image.kind == ResourceKind::StorageImageUint;
const auto format = (descriptor.dwords[1] >> 20u) & 0x1ffu;
const bool raw_sint_storage =
storage && format == Prospero::GpuEnumValue(Prospero::BufferFormat::k32SInt) &&
!image.read && !image.atomic;
const bool uint_descriptor = Prospero::IsUintTextureFormat(format) || raw_sint_storage;
const auto uint_program = image.kind == ResourceKind::ImageUint ||
image.kind == ResourceKind::StorageImageUint;
if (uint_descriptor != uint_program && !(image.atomic && uint_program)) {
if (error != nullptr) {
*error =
@@ -360,7 +405,9 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
const auto& descriptor = snapshot.images[i];
auto& image = next.images[i];
if (NullImageDescriptor(descriptor)) {
image.dimension = Decoder::ImageDimension::Dim2D;
image.dimension = Decoder::ImageDimension::Dim2D;
image.cube = false;
image.mip_levels = 1;
switch (image.kind) {
case ResourceKind::ImageUint: image.kind = ResourceKind::Image; break;
case ResourceKind::StorageImageUint:
@@ -372,6 +419,16 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
}
continue;
}
if (image.mip_mode == ImageMipMode::DynamicStorage) {
if (!DescriptorMipRange(descriptor, image.mip_levels)) {
if (error != nullptr) {
*error = fmt::format("image descriptor {} has invalid storage mip range", i);
}
return false;
}
} else {
image.mip_levels = 1;
}
const auto descriptor_dimension = DescriptorDimension(descriptor, image.dimension);
if (descriptor_dimension == Decoder::ImageDimension::Unknown) {
if (error != nullptr) {
@@ -386,11 +443,19 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
return false;
}
image.dimension = descriptor_dimension;
image.cube = DescriptorIsCube(descriptor);
if (image.kind == ResourceKind::StorageImage ||
image.kind == ResourceKind::StorageImageUint) {
image.storage_swizzle = DescriptorImageSwizzle(descriptor);
}
if (Prospero::IsUintTextureFormat((descriptor.dwords[1] >> 20u) & 0x1ffu)) {
const auto format = (descriptor.dwords[1] >> 20u) & 0x1ffu;
const bool storage = image.kind == ResourceKind::StorageImage ||
image.kind == ResourceKind::StorageImageUint;
const bool raw_sint_storage =
storage && format == Prospero::GpuEnumValue(Prospero::BufferFormat::k32SInt) &&
!image.read && !image.atomic;
const bool uint_image = Prospero::IsUintTextureFormat(format) || raw_sint_storage;
if (uint_image) {
switch (image.kind) {
case ResourceKind::Image: image.kind = ResourceKind::ImageUint; break;
case ResourceKind::StorageImage: image.kind = ResourceKind::StorageImageUint; break;
@@ -402,6 +467,7 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
std::reference_wrapper<Instruction> inst;
ResourceKind kind;
Decoder::ImageDimension dimension;
bool cube;
};
std::vector<ImagePatch> patches;
for (auto& block: program.blocks) {
@@ -420,13 +486,14 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
return false;
}
const auto& image = next.images[inst.memory.resource];
patches.push_back({std::ref(inst), image.kind, image.dimension});
patches.push_back({std::ref(inst), image.kind, image.dimension, image.cube});
}
}
program.info = std::move(next);
for (const auto& patch: patches) {
patch.inst.get().memory.kind = patch.kind;
patch.inst.get().memory.image_dimension = patch.dimension;
patch.inst.get().memory.image_cube = patch.cube;
}
return true;
}

Some files were not shown because too many files have changed in this diff Show More