Compare commits

...
Author SHA1 Message Date
nmzik d8a4c83cc7 format src and tests with clang-format 2026-07-31 11:36:12 +02:00
nmzik 68be13345a fix(shader): support multisampled depth image loads 2026-07-31 11:36:12 +02:00
nmzik 167da0abe0 shader: fix readlane/writelane for inactive host lanes 2026-07-31 11:36:12 +02:00
nmzik 48c31d61ee Implement VideoDec2 2026-07-31 11:36:12 +02:00
nmzik 212282d693 fix(shader): preserve packed UINT16 MRT exports 2026-07-31 11:36:12 +02:00
nmzik 6bca35d1f5 renderer: broaden compatibility 2026-07-31 11:36:12 +02:00
M. AbdullahandGitHub e4ad5fc988 docs: add macOS build and run instructions (#137)
The README had macOS badges and an experimental-support note but no build,
run, or system-requirement information for the platform. Document the
Rosetta 2 / MoltenVK setup, the x86-64 configure invocation, the Qt
universal-build requirement, MoltenVK installation and signing, and the
SDL_VULKAN_LIBRARY variable needed at run time.
2026-07-31 05:32:29 +02:00
nmzik c0d3d261ea add TextToSpeech2 stubs 2026-07-31 04:05:50 +02:00
3b75a5659a shader: specialize cube image descriptors (#134)
* shader: specialize cube image descriptors

Track whether image descriptors refer to cube maps during resource specialization, and apply the coordinate offset conversion when sampling cube maps as 2D image arrays in SPIR-V emission.

* shader: fix cube array coordinate lowering

---------

Co-authored-by: nmzik <Nmzik@mail.ru>
2026-07-31 03:59:10 +02:00
M. AbdullahandGitHub d475387171 macOS: enable guest signal dispatch on the target thread (#136)
macos: enable guest signal dispatch on the target thread

The POSIX signal-dispatch path (pthread_kill based, added with the Linux
port) was compiled out on macOS, leaving KernelRaiseException to run the
guest handler on the calling thread. IL2CPP's garbage collector raises its
stop-the-world signal at every managed thread and each handler parks its
own thread until resume, so the collector parked itself and every Unity
title froze on the first collection.

Enable the same delivery path on macOS:
- translate between the Darwin mcontext (uc_mcontext->__ss) and the guest
  ucontext in CreateSignalUcontextFromHost/ApplySignalUcontextToHost
- use SIGUSR1 as the host dispatch signal (macOS has no realtime signals)
- block the dispatch signal inside the host fault handler so a suspend
  request cannot preempt fault resolution between the protection fix and
  the retry

Windows and Linux are unchanged.
2026-07-31 03:47:04 +02:00
114 changed files with 21545 additions and 21613 deletions
+54 -5
View File
@@ -27,8 +27,9 @@ Development is focused on compatibility and boot reliability.
Windows is the primary platform and receives the most testing. Linux builds and runs; see
[Building on Linux](#building-on-linux).
macOS support is experimental. Compatibility with the same games on Windows and macOS has not yet
been tested.
macOS support is experimental. The emulator is built for x86-64 and runs on Apple Silicon under
Rosetta 2, with Vulkan provided by MoltenVK. A small number of titles have been verified in-game
on Apple Silicon hardware; see [Building on macOS](#building-on-macos).
## Bugs and Issues
@@ -114,9 +115,10 @@ the Vulkan/SPIR-V validation rules.
### System requirements
- Windows 10 version 1803, or a current Linux distribution
- A 64-bit x86 processor
- A Vulkan 1.3-capable GPU with current drivers
- Windows 10 version 1803, a current Linux distribution, or macOS on Apple Silicon
- A 64-bit x86 processor (on macOS, an Apple Silicon processor with Rosetta 2)
- A Vulkan 1.3-capable GPU with current drivers (on macOS, Vulkan is provided by the bundled
MoltenVK)
### Build requirements (Windows)
@@ -188,6 +190,45 @@ time.
Note that the CMake source root is `src`, not the repository root.
### Building on macOS
macOS builds target x86-64 and run under Rosetta 2 on Apple Silicon, so the PS5's x86-64 game
code executes through the same translation layer as the emulator itself. Prebuilt archives are
attached to releases; the steps below are for building from source.
Requirements:
- An Apple Silicon Mac with Rosetta 2 installed (`softwareupdate --install-rosetta`)
- Xcode (or the Command Line Tools)
- Homebrew packages: `brew install cmake ninja glslang`
- Qt 6 (Concurrent, Network, Widgets) with x86-64 support. The official Qt installation is
universal and works; Homebrew's Qt is arm64-only and will not link
```bash
git submodule update --init --recursive
cmake -S src -B _Build/macos -G Ninja -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_OSX_ARCHITECTURES=x86_64 \
-DCMAKE_C_COMPILER=clang -DCMAKE_CXX_COMPILER=clang++ \
-DCMAKE_PREFIX_PATH="$Qt6_DIR"
cmake --build _Build/macos --target launcher --parallel
cmake --install _Build/macos --prefix _Build/macos/install
```
The build re-signs `kyty_emulator` with the JIT entitlements it needs to execute translated
guest code; no manual signing step is required.
Vulkan comes from MoltenVK. Download `MoltenVK-macos.tar` from the
[MoltenVK releases](https://github.com/KhronosGroup/MoltenVK/releases), then copy
`MoltenVK/dynamic/dylib/macOS/libMoltenVK.dylib` next to `kyty_emulator` and ad-hoc sign it:
```bash
codesign --force --sign - _Build/macos/install/libMoltenVK.dylib
```
Release archives already include a signed `libMoltenVK.dylib`.
### Visual Studio Code
A ready-made Visual Studio Code setup is included in [`.vscode`](.vscode). It configures CMake
@@ -233,6 +274,14 @@ The emulator can also be started directly with a legally obtained game directory
./_Build/linux/install/kyty_emulator --game "/games/ExampleGame"
```
On macOS, point SDL at the MoltenVK library explicitly; the hardened runtime prevents it from
being picked up from the executable's directory:
```bash
cd _Build/macos/install
SDL_VULKAN_LIBRARY="$PWD/libMoltenVK.dylib" ./kyty_emulator --game "/games/ExampleGame"
```
Run `kyty_emulator --help` to see the available graphics, logging, validation, profiling, and
debugging options.
+2
View File
@@ -329,6 +329,7 @@ add_kyty_full_emulator_test(shader_cfg_tests ../tests/shaderCfgTests.cpp)
add_executable(scalar_provenance_tests EXCLUDE_FROM_ALL
../tests/ScalarProvenanceTests.cpp
graphics/host_gpu/hostMemory.cpp
graphics/shader/recompiler/ir/ReadLaneElimination.cpp
graphics/shader/recompiler/ir/ScalarProvenance.cpp
graphics/shader/recompiler/ir/SrtWalker.cpp
)
@@ -441,6 +442,7 @@ if(NOT KYTY_CLANG_CL)
endif()
if(BUILD_TESTING)
add_test(NAME scalar_provenance COMMAND $<TARGET_FILE:scalar_provenance_tests>)
add_test(NAME image_page_table COMMAND $<TARGET_FILE:image_page_table_tests>)
add_test(NAME memory_tracker COMMAND $<TARGET_FILE:memory_tracker_tests>)
add_test(NAME page_manager COMMAND $<TARGET_FILE:page_manager_tests>)
+10 -6
View File
@@ -175,9 +175,9 @@ static void SignalHandler(int sig, siginfo_t* si, void* uctx) {
}
g_in_exception_filter = true;
auto* uc = static_cast<ucontext_t*>(uctx);
const auto* mc = uc->uc_mcontext;
const auto& ss = mc->__ss;
auto* uc = static_cast<ucontext_t*>(uctx);
const auto* mc = uc->uc_mcontext;
const auto& ss = mc->__ss;
ExceptionInfo info {};
info.exception_address = ss.__rip;
@@ -214,7 +214,7 @@ static void SignalHandler(int sig, siginfo_t* si, void* uctx) {
FailFast("host exception callback is null");
}
const bool resolved = handler(info);
const bool resolved = handler(info);
g_in_exception_filter = false;
if (resolved) {
@@ -255,8 +255,8 @@ static void SignalHandler(int signal_number, siginfo_t* signal_info, void* nativ
info.native_context = context;
if (signal_number == SIGSEGV || signal_number == SIGBUS) {
info.type = ExceptionType::AccessViolation;
const auto error_code = static_cast<uint64_t>(gregs[REG_ERR]);
info.type = ExceptionType::AccessViolation;
const auto error_code = static_cast<uint64_t>(gregs[REG_ERR]);
if ((error_code & PAGE_FAULT_ERROR_INSTRUCTION) != 0) {
info.access_violation_type = AccessViolationType::Execute;
} else if ((error_code & PAGE_FAULT_ERROR_WRITE) != 0) {
@@ -324,6 +324,10 @@ bool InstallHandler(Handler handler) {
sa.sa_sigaction = SignalHandler;
sa.sa_flags = SA_SIGINFO;
sigemptyset(&sa.sa_mask);
// The guest signal-dispatch path (KernelRaiseException) interrupts threads with
// SIGUSR1; block it while a fault is being resolved so a stop-the-world request
// cannot preempt the handler between the protection fix and the retry.
sigaddset(&sa.sa_mask, SIGUSR1);
// macOS raises SIGBUS for protection faults on some paths and SIGSEGV on others;
// SIGILL covers instructions the host cannot execute (routed to the x64 emulator).
+7 -8
View File
@@ -19,10 +19,10 @@ class LeastRecentlyUsedCache {
public:
[[nodiscard]] size_t Insert(Object object, Tick tick) {
const auto id = Build();
const auto id = Build();
auto& item = m_items[id];
item.object = std::move(object);
item.tick = tick;
item.object = std::move(object);
item.tick = tick;
Attach(item);
return id;
}
@@ -49,8 +49,7 @@ public:
template <typename Function>
void ForEachItemBelow(Tick tick, Function&& function) {
constexpr bool ReturnsBool =
std::is_same_v<std::invoke_result_t<Function, Object>, bool>;
constexpr bool ReturnsBool = std::is_same_v<std::invoke_result_t<Function, Object>, bool>;
for (auto* item = m_first; item != nullptr;) {
if (item->tick > tick) {
return;
@@ -87,10 +86,10 @@ private:
m_last = &item;
return;
}
item.prev = m_last;
item.prev = m_last;
m_last->next = &item;
item.next = nullptr;
m_last = &item;
item.next = nullptr;
m_last = &item;
}
void Detach(Item& item) {
+3 -4
View File
@@ -31,10 +31,9 @@ static bool OnOwnStack() {
if (pthread_getattr_np(pthread_self(), &attr) != 0) {
return false;
}
void* base = nullptr;
size_t size = 0;
const bool ok =
pthread_attr_getstack(&attr, &base, &size) == 0 && base != nullptr && size != 0;
void* base = nullptr;
size_t size = 0;
const bool ok = pthread_attr_getstack(&attr, &base, &size) == 0 && base != nullptr && size != 0;
pthread_attr_destroy(&attr);
if (!ok) {
return false;
+3 -5
View File
@@ -172,8 +172,7 @@ sys_file_t* SysFileCreate(const std::filesystem::path& file_name) {
return ret;
}
sys_file_t* SysFileOpenR(const std::filesystem::path& file_name,
sys_file_cache_type_t cache_type) {
sys_file_t* SysFileOpenR(const std::filesystem::path& file_name, sys_file_cache_type_t cache_type) {
auto* ret = new sys_file_t;
ret->type = SYS_FILE_FILE;
@@ -218,8 +217,7 @@ sys_file_t* SysFileCreate() {
return ret;
}
sys_file_t* SysFileOpenW(const std::filesystem::path& file_name,
sys_file_cache_type_t cache_type) {
sys_file_t* SysFileOpenW(const std::filesystem::path& file_name, sys_file_cache_type_t cache_type) {
auto* ret = new sys_file_t;
auto real_name = get_internal_name(file_name);
@@ -241,7 +239,7 @@ sys_file_t* SysFileOpenW(const std::filesystem::path& file_name,
}
sys_file_t* SysFileOpenRw(const std::filesystem::path& file_name,
sys_file_cache_type_t cache_type) {
sys_file_cache_type_t cache_type) {
auto* ret = new sys_file_t;
auto real_name = get_internal_name(file_name);
+13 -13
View File
@@ -136,8 +136,8 @@ static void* map_anonymous(uintptr_t addr, size_t size, int protect, int flags)
break;
}
const auto hint = (top - step) & ~(LOW_ARENA_GRAIN - 1);
void* ptr = mmap(reinterpret_cast<void*>(hint), size, protect,
flags | MAP_FIXED_NOREPLACE, -1, 0); // NOLINT
void* ptr = mmap(reinterpret_cast<void*>(hint), size, protect, flags | MAP_FIXED_NOREPLACE,
-1, 0); // NOLINT
if (ptr != MAP_FAILED) {
return ptr;
}
@@ -161,8 +161,8 @@ uint64_t SysVirtualAlloc(uint64_t address, uint64_t size, VirtualMemory::Mode mo
if (ptr != MAP_FAILED) {
pthread_mutex_lock(&g_virtual_mutex);
record_alloc(ret_addr, size);
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
for (uintptr_t page = page_start; page <= page_end; page++) {
(*g_protects)[page] = protect;
}
@@ -194,8 +194,8 @@ uint64_t SysVirtualAllocAligned(uint64_t address, uint64_t size, VirtualMemory::
if (ptr != MAP_FAILED && ((ret_addr & (alignment - 1)) != 0)) {
munmap(ptr, size);
ptr = map_anonymous(addr, size + alignment, protect,
MAP_PRIVATE | MAP_ANON | MAP_NORESERVE);
ptr =
map_anonymous(addr, size + alignment, protect, MAP_PRIVATE | MAP_ANON | MAP_NORESERVE);
ret_addr = reinterpret_cast<uintptr_t>(ptr);
if (ptr != MAP_FAILED) {
#if defined(__APPLE__)
@@ -251,8 +251,8 @@ uint64_t SysVirtualAllocAligned(uint64_t address, uint64_t size, VirtualMemory::
pthread_mutex_lock(&g_virtual_mutex);
record_alloc(ret_addr, size);
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
for (uintptr_t page = page_start; page <= page_end; page++) {
(*g_protects)[page] = protect;
}
@@ -266,9 +266,9 @@ uint64_t SysVirtualAllocAligned(uint64_t address, uint64_t size, VirtualMemory::
// the first mapped region at or above `region_addr`; if it begins before the end of the
// requested range, the range overlaps an existing mapping.
static bool is_mapped(void* ptr, size_t length) {
auto query_addr = reinterpret_cast<mach_vm_address_t>(ptr);
mach_vm_address_t region_addr = query_addr;
mach_vm_size_t region_size = 0;
auto query_addr = reinterpret_cast<mach_vm_address_t>(ptr);
mach_vm_address_t region_addr = query_addr;
mach_vm_size_t region_size = 0;
vm_region_basic_info_data_64_t info {};
mach_msg_type_number_t count = VM_REGION_BASIC_INFO_COUNT_64;
mach_port_t object_name = MACH_PORT_NULL;
@@ -337,8 +337,8 @@ bool SysVirtualAllocFixed(uint64_t address, uint64_t size, VirtualMemory::Mode m
if (ptr != MAP_FAILED) {
pthread_mutex_lock(&g_virtual_mutex);
record_alloc(ret_addr, size);
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
uintptr_t page_start = ret_addr >> 12u;
uintptr_t page_end = (ret_addr + size - 1) >> 12u;
for (uintptr_t page = page_start; page <= page_end; page++) {
(*g_protects)[page] = protect;
}
+1 -1
View File
@@ -5,9 +5,9 @@
#include <algorithm>
#include <atomic>
#include <cerrno>
#include <chrono> // IWYU pragma: keep
#include <condition_variable> // IWYU pragma: keep
#include <cerrno>
#include <mutex>
#include <vector>
+2 -4
View File
@@ -11,7 +11,7 @@ template <typename Result, typename... Args>
class UniqueFunction {
class CallableBase {
public:
virtual ~CallableBase() = default;
virtual ~CallableBase() = default;
virtual Result Invoke(Args&&... args) = 0;
};
@@ -20,9 +20,7 @@ class UniqueFunction {
public:
explicit Callable(Function function): m_function(std::move(function)) {}
Result Invoke(Args&&... args) override {
return m_function(std::forward<Args>(args)...);
}
Result Invoke(Args&&... args) override { return m_function(std::forward<Args>(args)...); }
private:
Function m_function;
@@ -158,7 +158,7 @@ private:
void CheckBuffer() const { GetScheduler().CheckActive(); }
GpuResourceManager& GetGpuResources() const { return m_renderer.GetGpuResources(); }
RenderContext& m_renderer;
RenderContext& m_renderer;
HW::Context m_ctx;
HW::UserConfig m_ucfg;
HW::Shader m_sh_ctx;
@@ -170,9 +170,9 @@ private:
uint64_t m_dispatch_indirect_args_base_addr = 0;
uint32_t m_num_instances = 1;
uint32_t m_de_count = 0;
uint32_t m_ce_count = 0;
bool m_ce_complete = false;
uint32_t m_de_count = 0;
uint32_t m_ce_count = 0;
bool m_ce_complete = false;
bool m_readback_active = false;
uint32_t m_const_ram[0x3000] = {0};
+2
View File
@@ -374,6 +374,8 @@ enum class BufferFormat : uint32_t {
k32_32_32_32UInt = 75,
k32_32_32_32SInt = 76,
k32_32_32_32Float = 77,
k8Srgb = 128,
k8_8Srgb = 129,
k8_8_8_8Srgb = 130,
k9_9_9_5Float = 132,
k5_6_5UNorm = 133,
+2
View File
@@ -57,6 +57,8 @@ constexpr FormatInfo kFormatInfo[] = {
{GpuEnumValue(BufferFormat::k32_32_32_32UInt), 16, 0, 16, true, true},
{GpuEnumValue(BufferFormat::k32_32_32_32SInt), 16, 0, 16, false, false},
{GpuEnumValue(BufferFormat::k32_32_32_32Float), 16, 0, 16, true, false},
{GpuEnumValue(BufferFormat::k8Srgb), 1, 0, 0, true, false},
{GpuEnumValue(BufferFormat::k8_8Srgb), 2, 0, 0, true, false},
{GpuEnumValue(BufferFormat::k8_8_8_8Srgb), 4, 0, 4, true, false},
{GpuEnumValue(BufferFormat::k9_9_9_5Float), 4, 0, 0, true, false},
{GpuEnumValue(BufferFormat::k5_6_5UNorm), 2, 0, 2, true, false},
+11 -13
View File
@@ -962,9 +962,8 @@ void CommandProcessor::DrawIndexOffset(uint32_t index_offset, uint32_t index_cou
auto* index_addr = reinterpret_cast<const void*>(
m_index_base_addr + static_cast<uint64_t>(index_offset) * index_size);
m_renderer.GetRenderExecutor().DrawIndex(m_submit_id, CurrentBuffer(),
m_index_type_and_size, index_count, index_addr,
flags, 1, m_num_instances);
m_renderer.GetRenderExecutor().DrawIndex(m_submit_id, CurrentBuffer(), m_index_type_and_size,
index_count, index_addr, flags, 1, m_num_instances);
}
void CommandProcessor::DrawIndirect(uint32_t data_offset, uint32_t draw_initiator, bool indexed) {
@@ -1190,8 +1189,8 @@ void CommandProcessor::DispatchDirect(uint32_t thread_group_x, uint32_t thread_g
}
}
m_renderer.GetRenderExecutor().DispatchDirect(
m_submit_id, CurrentBuffer(), thread_group_x, thread_group_y, thread_group_z, mode);
m_renderer.GetRenderExecutor().DispatchDirect(m_submit_id, CurrentBuffer(), thread_group_x,
thread_group_y, thread_group_z, mode);
}
constexpr uint32_t DispatchInitiatorUseThreadDimensions = 1u << 5u;
@@ -1237,16 +1236,16 @@ void CommandProcessor::DrawIndexAuto(uint32_t index_count, uint32_t flags,
uint32_t first_vertex, uint32_t first_instance) {
CheckBuffer();
m_renderer.GetRenderExecutor().DrawAuto(
m_submit_id, CurrentBuffer(), index_count, flags, render_target_slice_offset,
instance_count, first_vertex, first_instance);
m_renderer.GetRenderExecutor().DrawAuto(m_submit_id, CurrentBuffer(), index_count, flags,
render_target_slice_offset, instance_count,
first_vertex, first_instance);
}
void CommandProcessor::WaitFlipDone(uint32_t video_out_handle, uint32_t display_buffer_index) {
BufferFlush();
m_renderer.GetVideoOut().WaitFlipDone(static_cast<int>(video_out_handle),
static_cast<int>(display_buffer_index));
static_cast<int>(display_buffer_index));
}
template <typename T>
@@ -1317,8 +1316,8 @@ void CommandProcessor::WriteAtEndOfPipe(uint32_t cache_policy, uint32_t event_wr
if (eop_event_type == 0x2f && cache_action == 0x00 && event_index == 0x06) {
auto* dst = static_cast<uint32_t*>(dst_gpu_addr);
SynchronizeGpu();
Sync::ReadGds(m_renderer.GetBufferCache().GetGdsBuffer(), dst,
value & 0xffffu, value >> 16u);
Sync::ReadGds(m_renderer.GetBufferCache().GetGdsBuffer(), dst, value & 0xffffu,
value >> 16u);
Sync::WriteAtEndOfPipeGds32(m_submit_id, CurrentBuffer(), dst, value & 0xffffu,
value >> 16u);
return;
@@ -1486,8 +1485,7 @@ void CommandProcessor::EmitGlobalBarrier() {
barrier.srcStageMask = vk::PipelineStageFlagBits2::eAllCommands;
barrier.srcAccessMask = vk::AccessFlagBits2::eMemoryWrite;
barrier.dstStageMask = vk::PipelineStageFlagBits2::eAllCommands;
barrier.dstAccessMask =
vk::AccessFlagBits2::eMemoryRead | vk::AccessFlagBits2::eMemoryWrite;
barrier.dstAccessMask = vk::AccessFlagBits2::eMemoryRead | vk::AccessFlagBits2::eMemoryWrite;
vk::DependencyInfo dependency {};
dependency.memoryBarrierCount = 1;
+11 -11
View File
@@ -65,41 +65,41 @@ struct TileVolumeLayout {
};
bool TileGetBlockLayout(TileBlockFamily family, uint32_t bytes_per_element,
TileBlockLayout& layout);
TileBlockLayout& layout);
bool TileGetBlockOffset(const TileBlockLayout& layout, uint32_t x, uint32_t y, uint32_t z,
uint32_t& byte_offset);
uint32_t& byte_offset);
bool TileGetBlockXor(const TileBlockLayout& layout, uint32_t block_x, uint32_t block_y,
uint32_t& byte_offset);
uint32_t& byte_offset);
bool TileGetBlockXor(const TileBlockLayout& layout, uint32_t block_x, uint32_t block_y,
uint32_t block_z, uint32_t& byte_offset);
uint32_t block_z, uint32_t& byte_offset);
bool TileIsStandard256BTextureSupported(uint32_t format);
bool TileIsStandard4KBTextureSupported(uint32_t format);
bool TileIsStandard64KBTextureSupported(uint32_t format);
bool TileGetTextureVolumeLayout(uint32_t format, uint32_t width, uint32_t height, uint32_t depth,
uint32_t levels, uint32_t tile, TileVolumeLayout& layout);
uint32_t levels, uint32_t tile, TileVolumeLayout& layout);
bool TileGetHtileSize(uint32_t width, uint32_t height, TileSizeAlign& htile_size);
bool TileGetDepthSize(uint32_t width, uint32_t height, uint32_t pitch, uint32_t z_format,
uint32_t stencil_format, bool htile, TileSizeAlign& stencil_size,
TileSizeAlign& htile_size, TileSizeAlign& depth_size,
uint32_t stencil_format, bool htile, TileSizeAlign& stencil_size,
TileSizeAlign& htile_size, TileSizeAlign& depth_size,
uint32_t num_fragments_log2 = 0);
uint32_t TileGetRenderTargetPitch(uint32_t width, uint32_t bytes_per_element,
uint32_t num_fragments_log2 = 0);
uint32_t TileGetDepthPitch(uint32_t width, uint32_t bytes_per_element,
uint32_t num_fragments_log2 = 0);
bool TileGetRenderTargetSize(uint32_t width, uint32_t height, uint32_t pitch,
uint32_t bytes_per_element, TileSizeAlign& total_size,
uint32_t bytes_per_element, TileSizeAlign& total_size,
uint32_t num_fragments_log2 = 0);
bool TileGetRenderTargetMipLayout(uint32_t width, uint32_t height, uint32_t pitch,
uint32_t bytes_per_element, uint32_t levels,
TileSizeAlign& total_size, TileSizeOffset* level_sizes,
uint32_t bytes_per_element, uint32_t levels,
TileSizeAlign& total_size, TileSizeOffset* level_sizes,
TilePaddedSize* padded_size);
void TileGetTextureSize(uint32_t format, uint32_t width, uint32_t height, uint32_t pitch,
uint32_t levels, uint32_t tile, TileSizeAlign* total_size,
TileSizeOffset* level_sizes, TilePaddedSize* padded_size);
void TileGetTextureTotalSize(uint32_t format, uint32_t width, uint32_t height, uint32_t depth,
uint32_t pitch, uint32_t levels, uint32_t tile, bool volume_texture,
TileSizeAlign& total_size);
TileSizeAlign& total_size);
uint32_t TileGetTexturePitch(uint32_t format, uint32_t width, uint32_t levels, uint32_t tile);
} // namespace Libs::Graphics
+12 -12
View File
@@ -60,19 +60,19 @@ struct VulkanImage {
VulkanImage() = default;
KYTY_CLASS_NO_COPY(VulkanImage);
vk::Format format = vk::Format::eUndefined;
vk::ImageType image_type = vk::ImageType::e2D;
vk::Extent3D extent = {1, 1, 1};
uint32_t guest_pitch = 0;
uint32_t layers = 1;
uint32_t mip_levels = 1;
uint32_t samples = 1;
vk::ImageUsageFlags usage = {};
vk::ImageCreateFlags flags = {};
vk::Image image = nullptr;
VulkanImageState state;
vk::Format format = vk::Format::eUndefined;
vk::ImageType image_type = vk::ImageType::e2D;
vk::Extent3D extent = {1, 1, 1};
uint32_t guest_pitch = 0;
uint32_t layers = 1;
uint32_t mip_levels = 1;
uint32_t samples = 1;
vk::ImageUsageFlags usage = {};
vk::ImageCreateFlags flags = {};
vk::Image image = nullptr;
VulkanImageState state;
std::vector<VulkanImageState> subresource_states;
Graphics::VulkanMemory memory;
Graphics::VulkanMemory memory;
};
struct VulkanBuffer {
+1 -1
View File
@@ -30,7 +30,7 @@ bool IsAccessible(DWORD protect, HostMemoryAccess access) {
} // namespace
bool HostMemoryQueryRange(uint64_t addr, uint64_t requested_size, HostMemoryAccess access,
uint64_t& accessible_size) {
uint64_t& accessible_size) {
accessible_size = 0;
if (addr == 0 || requested_size == 0) {
return false;
+1 -1
View File
@@ -8,7 +8,7 @@ namespace Libs::Graphics {
enum class HostMemoryAccess { Read, Mapped };
bool HostMemoryQueryRange(uint64_t addr, uint64_t requested_size, HostMemoryAccess access,
uint64_t& accessible_size);
uint64_t& accessible_size);
bool HostMemoryQueryReadable(uint64_t addr, uint64_t requested_size, uint64_t& readable_size);
bool HostMemoryIsReadable(uint64_t addr);
bool HostMemoryRangeIsReadable(uint64_t addr, uint64_t size);
+10 -12
View File
@@ -135,8 +135,8 @@ void Buffer::Write(uint64_t offset, const void* source, uint64_t size) {
void Buffer::Flush(uint64_t offset, uint64_t size) {
EXIT_IF(m_mapped.empty() || offset > m_size || size > m_size - offset);
if (!m_is_coherent && size != 0) {
const auto result = vmaFlushAllocation(m_graphics->allocator, m_buffer->memory.allocation,
offset, size);
const auto result =
vmaFlushAllocation(m_graphics->allocator, m_buffer->memory.allocation, offset, size);
EXIT_NOT_IMPLEMENTED(static_cast<vk::Result>(result) != vk::Result::eSuccess);
}
}
@@ -144,8 +144,8 @@ void Buffer::Flush(uint64_t offset, uint64_t size) {
vk::BufferMemoryBarrier Buffer::Barrier(uint64_t offset, uint64_t size, vk::AccessFlags source,
vk::AccessFlags destination) const {
if (Handle() == nullptr || size == 0 || offset > m_size || size > m_size - offset) {
EXIT("Buffer: invalid DMA barrier, handle=%p offset=0x%016" PRIx64
" size=0x%016" PRIx64 " capacity=0x%016" PRIx64 "\n",
EXIT("Buffer: invalid DMA barrier, handle=%p offset=0x%016" PRIx64 " size=0x%016" PRIx64
" capacity=0x%016" PRIx64 "\n",
static_cast<const void*>(Handle()), offset, size, m_size);
}
vk::BufferMemoryBarrier barrier {};
@@ -175,10 +175,9 @@ void Buffer::CopyFrom(CommandBuffer& command, const Buffer& source, uint64_t sou
command.EndRendering();
const vk::BufferMemoryBarrier before[] = {
source.Barrier(source_offset, size, source_before, vk::AccessFlagBits::eTransferRead),
Barrier(destination_offset, size, destination_before,
vk::AccessFlagBits::eTransferWrite),
Barrier(destination_offset, size, destination_before, vk::AccessFlagBits::eTransferWrite),
};
const auto host_access = vk::AccessFlagBits::eHostRead | vk::AccessFlagBits::eHostWrite;
const auto host_access = vk::AccessFlagBits::eHostRead | vk::AccessFlagBits::eHostWrite;
auto before_stage = vk::PipelineStageFlags {vk::PipelineStageFlagBits::eAllCommands};
if (static_cast<bool>((source_before | destination_before) & host_access)) {
before_stage |= vk::PipelineStageFlagBits::eHost;
@@ -214,9 +213,8 @@ void Buffer::Fill(uint64_t offset, uint64_t size, uint32_t value) {
vk::PipelineStageFlagBits::eTransfer, vk::DependencyFlagBits::eByRegion,
0, nullptr, 1, &before, 0, nullptr);
native.fillBuffer(Handle(), offset, size, value);
const auto after =
Barrier(offset, size, vk::AccessFlagBits::eTransferWrite,
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite);
const auto after = Barrier(offset, size, vk::AccessFlagBits::eTransferWrite,
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite);
native.pipelineBarrier(vk::PipelineStageFlagBits::eTransfer,
vk::PipelineStageFlagBits::eAllCommands,
vk::DependencyFlagBits::eByRegion, 0, nullptr, 1, &after, 0, nullptr);
@@ -250,8 +248,8 @@ std::pair<uint8_t*, uint64_t> StreamBuffer::Map(uint64_t size, uint64_t alignmen
if (Mapped().empty()) {
return {nullptr, 0};
}
uint64_t mapped_size = size;
const auto atom = Graphics().physical_device_properties.limits.nonCoherentAtomSize;
uint64_t mapped_size = size;
const auto atom = Graphics().physical_device_properties.limits.nonCoherentAtomSize;
if (!NormalizeReservation(IsCoherent(), atom, mapped_size, alignment)) {
return {nullptr, 0};
}
+14 -15
View File
@@ -54,16 +54,15 @@ public:
[[nodiscard]] bool IsInBounds(uint64_t address, uint64_t size) const noexcept;
void Write(uint64_t offset, const void* source, uint64_t size);
void Flush(uint64_t offset, uint64_t size);
void CopyFrom(
CommandBuffer& command, const Buffer& source, uint64_t source_offset,
uint64_t destination_offset, uint64_t size,
vk::AccessFlags source_before = vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_before =
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags source_after =
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_after =
vk::AccessFlagBits::eMemoryRead | vk::AccessFlagBits::eMemoryWrite);
void CopyFrom(CommandBuffer& command, const Buffer& source, uint64_t source_offset,
uint64_t destination_offset, uint64_t size,
vk::AccessFlags source_before = vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_before = vk::AccessFlagBits::eMemoryRead |
vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags source_after = vk::AccessFlagBits::eMemoryRead |
vk::AccessFlagBits::eMemoryWrite,
vk::AccessFlags destination_after = vk::AccessFlagBits::eMemoryRead |
vk::AccessFlagBits::eMemoryWrite);
void Fill(uint64_t offset, uint64_t size, uint32_t value);
protected:
@@ -107,13 +106,13 @@ private:
uint64_t upper_bound = 0;
};
void ReserveWatches(std::vector<Watch>& watches, size_t grow_size);
void ReserveWatches(std::vector<Watch>& watches, size_t grow_size);
[[nodiscard]] static bool NormalizeReservation(bool coherent, uint64_t atom, uint64_t& size,
uint64_t& alignment);
[[nodiscard]] bool WaitPendingOperations(const std::vector<Watch>& watches,
std::optional<size_t> invalidation_mark,
uint64_t requested_upper_bound, bool allow_wait,
size_t& wait_cursor, uint64_t& wait_bound);
[[nodiscard]] bool WaitPendingOperations(const std::vector<Watch>& watches,
std::optional<size_t> invalidation_mark,
uint64_t requested_upper_bound, bool allow_wait,
size_t& wait_cursor, uint64_t& wait_bound);
uint64_t m_offset = 0;
uint64_t m_mapped_size = 0;
@@ -7,8 +7,8 @@
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/guest_gpu/tile.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
@@ -23,10 +23,10 @@ static std::atomic<uint32_t> g_render_color_log_count = 0;
// NOLINTNEXTLINE(readability-function-cognitive-complexity)
void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandBuffer& buffer,
RenderColorInfo& r,
uint32_t render_target_slice_offset,
uint32_t render_target_slot, bool ignore_target_mask,
bool exact_format) {
RenderColorInfo& r,
uint32_t render_target_slice_offset,
uint32_t render_target_slot, bool ignore_target_mask,
bool exact_format) {
KYTY_PROFILER_FUNCTION();
const auto& hw = buffer.GetRegisters();
@@ -79,10 +79,8 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
const auto view = ResolveTargetViewInfo(
rt.view.base_array_slice_index, rt.view.last_array_slice_index, render_target_slice_offset);
switch (view.type) {
case TargetViewType::Image2D: break;
case TargetViewType::Image2DArray:
EXIT("layered render-target views are unsupported: base=%u count=%u\n", view.base_layer,
view.layer_count);
case TargetViewType::Image2D:
case TargetViewType::Image2DArray: break;
case TargetViewType::Unsupported:
EXIT("invalid render-target view: base=%u last=%u draw_offset=%u\n",
rt.view.base_array_slice_index, rt.view.last_array_slice_index,
@@ -241,12 +239,12 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
}
TextureCache::ImageDesc desc {};
desc.type = TextureCache::BindingType::RenderTarget;
desc.info.data = {rt.base.addr, backing_size};
desc.info.pixel_format = target_format.format;
desc.info.guest_format = ImageOps::RenderTargetTransferFormat(bytes_per_element);
desc.info.type = Prospero::ImageType::kColor2D;
desc.info.extent = {width, height, 1};
desc.type = TextureCache::BindingType::RenderTarget;
desc.info.data = {rt.base.addr, backing_size};
desc.info.pixel_format = target_format.format;
desc.info.guest_format = ImageOps::RenderTargetTransferFormat(bytes_per_element);
desc.info.type = Prospero::ImageType::kColor2D;
desc.info.extent = {width, height, 1};
desc.info.resources = {levels, view.image_layers};
desc.info.pitch = pitch;
desc.info.bytes_per_block = bytes_per_element;
@@ -275,20 +273,20 @@ void RenderExecutor::ResolveRenderColorTarget(uint64_t submit_id, RenderCommandB
desc.view_info.base_layer = view.base_layer;
desc.view_info.layer_count = view.layer_count;
desc.view_info.usage = vk::ImageUsageFlagBits::eColorAttachment;
auto& texture_cache = m_context.GetTextureCache();
r.desc = std::move(desc);
r.image_id = texture_cache.FindImage(r.desc, exact_format);
r.type = RenderColorType::RenderTexture;
r.base_addr = rt.base.addr;
r.image_view = nullptr;
r.format = r.desc.view_info.format;
r.extent = view_extent;
r.base_mip_level = rt.view.current_mip_level;
r.buffer_size = backing_size;
r.samples = samples;
r.export_mapping = target_format.export_mapping;
r.color_clear_enable = false;
r.color_clear_value = {};
auto& texture_cache = m_context.GetTextureCache();
r.desc = std::move(desc);
r.image_id = texture_cache.FindImage(r.desc, exact_format);
r.type = RenderColorType::RenderTexture;
r.base_addr = rt.base.addr;
r.image_view = nullptr;
r.format = r.desc.view_info.format;
r.extent = view_extent;
r.base_mip_level = rt.view.current_mip_level;
r.buffer_size = backing_size;
r.samples = samples;
r.export_mapping = target_format.export_mapping;
r.color_clear_enable = false;
r.color_clear_value = {};
BindRenderTarget(r.image_id);
}
@@ -2,8 +2,8 @@
#define EMULATOR_SRC_GRAPHICS_HOST_GPU_RENDERER_COLORRENDERTARGET_H_
#include "graphics/guest_gpu/gpu_defs.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include <cstdint>
@@ -44,13 +44,13 @@ CommandSlot* CommandScheduler::CommandPool::CreateSlot() {
allocate.commandPool = m_pool;
allocate.level = vk::CommandBufferLevel::ePrimary;
allocate.commandBufferCount = 1;
vk::CommandBuffer buffer = nullptr;
vk::CommandBuffer buffer = nullptr;
EXIT_IF(graphics.device.allocateCommandBuffers(&allocate, &buffer) != vk::Result::eSuccess);
vk::FenceCreateInfo fence_create {};
fence_create.sType = vk::StructureType::eFenceCreateInfo;
fence_create.flags = vk::FenceCreateFlagBits::eSignaled;
vk::Fence fence = nullptr;
vk::Fence fence = nullptr;
if (graphics.device.createFence(&fence_create, nullptr, &fence) != vk::Result::eSuccess) {
graphics.device.freeCommandBuffers(m_pool, 1, &buffer);
EXIT("failed to create command-buffer fence\n");
@@ -70,9 +70,9 @@ CommandSlot* CommandScheduler::CommandPool::Allocate(GraphicContext& graphics) {
Create(graphics);
}
EXIT_IF(m_graphics != &graphics);
auto found = std::ranges::find_if(m_slots, [](const auto& slot) { return !slot.busy; });
auto* slot = found != m_slots.end() ? &*found : CreateSlot();
slot->busy = true;
auto found = std::ranges::find_if(m_slots, [](const auto& slot) { return !slot.busy; });
auto* slot = found != m_slots.end() ? &*found : CreateSlot();
slot->busy = true;
slot->Reset();
return slot;
}
@@ -331,8 +331,7 @@ void CommandScheduler::WaitPriorityOperations(uint64_t tick) {
EXIT_IF(g_deferred_callback_scheduler == this);
std::unique_lock lock(m_operation_mutex);
m_operation_available.wait(lock, [this, tick] {
const bool active_before_or_at =
m_priority_active && m_priority_active_tick <= tick;
const bool active_before_or_at = m_priority_active && m_priority_active_tick <= tick;
const bool queued_before_or_at =
!m_priority_operations.empty() && m_priority_operations.front().tick <= tick;
return !active_before_or_at && !queued_before_or_at;
@@ -47,21 +47,21 @@ public:
void FinishCurrent();
// Deferred callbacks can observe an externally owned drain, but cannot initiate shutdown:
// the priority runner cannot join itself.
void Shutdown();
void Wait(uint64_t tick);
void PopPendingOperations();
void DrainPriorityOperations();
void WaitPriorityOperations(uint64_t tick);
void DeferOperation(Common::UniqueFunction<void>&& operation);
void DeferPriorityOperation(Common::UniqueFunction<void>&& operation);
void Shutdown();
void Wait(uint64_t tick);
void PopPendingOperations();
void DrainPriorityOperations();
void WaitPriorityOperations(uint64_t tick);
void DeferOperation(Common::UniqueFunction<void>&& operation);
void DeferPriorityOperation(Common::UniqueFunction<void>&& operation);
[[nodiscard]] static bool InDeferredOperation() noexcept;
[[nodiscard]] bool Active() const noexcept { return m_current >= 0; }
void CheckActive() const;
RenderCommandBuffer& Current() const;
[[nodiscard]] uint64_t CurrentTick() const noexcept { return m_master.CurrentTick(); }
[[nodiscard]] bool IsFree(uint64_t tick);
[[nodiscard]] RenderContext& Context() const noexcept { return m_context; }
[[nodiscard]] bool Active() const noexcept { return m_current >= 0; }
void CheckActive() const;
RenderCommandBuffer& Current() const;
[[nodiscard]] uint64_t CurrentTick() const noexcept { return m_master.CurrentTick(); }
[[nodiscard]] bool IsFree(uint64_t tick);
[[nodiscard]] RenderContext& Context() const noexcept { return m_context; }
[[nodiscard]] GraphicContext& Graphics() const noexcept { return m_graphics; }
private:
@@ -91,11 +91,11 @@ private:
uint64_t tick = 0;
};
void BindCurrent() const;
CommandBuffer& SubmitCurrent(SubmitInfo& submit);
void BeginNext();
void PriorityOperationsThread(std::stop_token stop);
void RunOperation(Common::UniqueFunction<void>&& operation);
void BindCurrent() const;
CommandBuffer& SubmitCurrent(SubmitInfo& submit);
void BeginNext();
void PriorityOperationsThread(std::stop_token stop);
void RunOperation(Common::UniqueFunction<void>&& operation);
[[nodiscard]] CommandSlot* AllocateCommandBuffer();
[[nodiscard]] uint64_t NextSubmitSequence() noexcept;
@@ -109,14 +109,14 @@ private:
std::mutex m_operation_mutex;
std::condition_variable m_operation_available;
std::jthread m_priority_thread;
bool m_priority_active = false;
bool m_priority_active = false;
uint64_t m_priority_active_tick = 0;
OperationState m_operation_state = OperationState::Open;
int m_current = -1;
bool m_recording = false;
HW::Context* m_registers = nullptr;
HW::UserConfig* m_user_config = nullptr;
HW::Shader* m_shaders = nullptr;
OperationState m_operation_state = OperationState::Open;
int m_current = -1;
bool m_recording = false;
HW::Context* m_registers = nullptr;
HW::UserConfig* m_user_config = nullptr;
HW::Shader* m_shaders = nullptr;
std::atomic<uint64_t> m_submit_sequence = 0;
friend class CommandBuffer;
+13 -13
View File
@@ -8,8 +8,8 @@
#include "graphics/host_gpu/renderer/colorRenderTarget.h"
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/depthRenderTarget.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/image/imageView.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/vma.h"
@@ -270,30 +270,30 @@ void CommandBuffer::BeginRendering(const RenderState& state) const {
colors[i].sType = vk::StructureType::eRenderingAttachmentInfo;
colors[i].imageView = attachment.image_view;
colors[i].imageLayout = attachment.image_layout;
colors[i].loadOp = attachment.is_clear ? vk::AttachmentLoadOp::eClear
: vk::AttachmentLoadOp::eLoad;
colors[i].storeOp = vk::AttachmentStoreOp::eStore;
colors[i].clearValue.color.uint32 = attachment.clear_value;
colors[i].loadOp =
attachment.is_clear ? vk::AttachmentLoadOp::eClear : vk::AttachmentLoadOp::eLoad;
colors[i].storeOp = vk::AttachmentStoreOp::eStore;
colors[i].clearValue.color.uint32 = attachment.clear_value;
}
const auto& depth_stencil = state.depth_stencil_attachment;
const auto& depth_stencil = state.depth_stencil_attachment;
vk::RenderingAttachmentInfo depth {};
depth.sType = vk::StructureType::eRenderingAttachmentInfo;
depth.imageView = depth_stencil.image_view;
depth.imageLayout = depth_stencil.image_layout;
depth.loadOp = depth_stencil.depth_clear ? vk::AttachmentLoadOp::eClear
: vk::AttachmentLoadOp::eLoad;
depth.storeOp = vk::AttachmentStoreOp::eStore;
depth.loadOp =
depth_stencil.depth_clear ? vk::AttachmentLoadOp::eClear : vk::AttachmentLoadOp::eLoad;
depth.storeOp = vk::AttachmentStoreOp::eStore;
depth.clearValue.depthStencil.depth = std::bit_cast<float>(depth_stencil.clear_value[0]);
vk::RenderingAttachmentInfo stencil {};
stencil.sType = vk::StructureType::eRenderingAttachmentInfo;
stencil.imageView = depth_stencil.image_view;
stencil.imageLayout = depth_stencil.image_layout;
stencil.loadOp = depth_stencil.stencil_clear ? vk::AttachmentLoadOp::eClear
: vk::AttachmentLoadOp::eLoad;
stencil.storeOp = vk::AttachmentStoreOp::eStore;
stencil.clearValue.depthStencil.stencil = depth_stencil.clear_value[1];
stencil.loadOp =
depth_stencil.stencil_clear ? vk::AttachmentLoadOp::eClear : vk::AttachmentLoadOp::eLoad;
stencil.storeOp = vk::AttachmentStoreOp::eStore;
stencil.clearValue.depthStencil.stencil = depth_stencil.clear_value[1];
vk::RenderingInfo rendering {};
rendering.sType = vk::StructureType::eRenderingInfo;
-8
View File
@@ -548,14 +548,6 @@ static void ZCheck(const HW::DepthRenderTarget& z) {
EXIT_NOT_IMPLEMENTED(z.htile_surface.prefetch_height != 0x00000000);
EXIT_NOT_IMPLEMENTED(z.htile_surface.dst_outside_zero_to_one != 0x00000000);
if (z.depth_view.slice_start != 0x00000000 || z.depth_view.slice_max != 0x00000000) {
static std::atomic<uint32_t> log_count {0};
if (log_count.fetch_add(1, std::memory_order_relaxed) < 16) {
LOGF("DepthTarget: temporary: ignoring PS5 array slice view start=0x%08" PRIx32
", max=0x%08" PRIx32 "\n",
z.depth_view.slice_start, z.depth_view.slice_max);
}
}
if (z.depth_view.current_mip_level != 0x00000000) {
static std::atomic<uint32_t> log_count {0};
if (log_count.fetch_add(1, std::memory_order_relaxed) < 16) {
@@ -10,10 +10,10 @@
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/guest_gpu/tile.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/image/imageView.h"
#include "graphics/host_gpu/renderer/image/textureCommon.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/vulkanCommon.h"
@@ -150,10 +150,8 @@ void RenderExecutor::ResolveRenderDepthTarget(uint64_t submit_id, RenderCommandB
has_stencil, has_htile, z.stencil_info.htile_stencil_disabled);
const auto view = ResolveTargetViewInfo(z.depth_view.slice_start, z.depth_view.slice_max);
switch (view.type) {
case TargetViewType::Image2D: break;
case TargetViewType::Image2DArray:
DepthFatal("layered depth views are unsupported: base=%u count=%u", view.base_layer,
view.layer_count);
case TargetViewType::Image2D:
case TargetViewType::Image2DArray: break;
case TargetViewType::Unsupported:
DepthFatal("invalid depth view: base=%u last=%u", z.depth_view.slice_start,
z.depth_view.slice_max);
@@ -2,9 +2,9 @@
#define EMULATOR_SRC_GRAPHICS_HOST_GPU_RENDERER_DEPTHRENDERTARGET_H_
#include "common/assert.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/renderer/image/imageView.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include <cstdint>
@@ -182,8 +182,8 @@ void BlitHelper::ReinterpretColorAsMsDepth(Image& source, Image& destination) {
auto command = command_buffer.Handle();
source.Transit(vk::ImageLayout::eShaderReadOnlyOptimal, vk::AccessFlagBits2::eShaderRead, {},
command);
destination.Transit(ColorToMsDepthLayout,
vk::AccessFlagBits2::eDepthStencilAttachmentWrite, {}, command);
destination.Transit(ColorToMsDepthLayout, vk::AccessFlagBits2::eDepthStencilAttachmentWrite, {},
command);
vk::RenderingAttachmentInfo depth_attachment {};
depth_attachment.sType = vk::StructureType::eRenderingAttachmentInfo;
@@ -19,10 +19,9 @@ struct GuestRange {
uint64_t address = 0;
uint64_t size = 0;
[[nodiscard]] constexpr bool Empty() const noexcept { return address == 0 || size == 0; }
[[nodiscard]] constexpr bool Valid() const noexcept {
return !Empty() && address < TRACKER_ADDRESS_SIZE &&
size <= TRACKER_ADDRESS_SIZE - address;
[[nodiscard]] constexpr bool Empty() const noexcept { return address == 0 || size == 0; }
[[nodiscard]] constexpr bool Valid() const noexcept {
return !Empty() && address < TRACKER_ADDRESS_SIZE && size <= TRACKER_ADDRESS_SIZE - address;
}
[[nodiscard]] constexpr uint64_t End() const noexcept { return address + size; }
auto operator<=>(const GuestRange&) const = default;
@@ -47,10 +46,10 @@ struct ImageSubresources {
};
struct ImageSubresourceRange {
uint32_t base_level = 0;
uint32_t level_count = 1;
uint32_t base_layer = 0;
uint32_t layer_count = 1;
uint32_t base_level = 0;
uint32_t level_count = 1;
uint32_t base_layer = 0;
uint32_t layer_count = 1;
auto operator<=>(const ImageSubresourceRange&) const = default;
};
@@ -67,10 +66,10 @@ struct ImageInfo {
GuestRange stencil;
ImageMetadataInfo metadata;
uint32_t htile_clear_mask = UINT32_MAX;
vk::Format pixel_format = vk::Format::eUndefined;
uint32_t guest_format = 0;
Prospero::ImageType type = Prospero::ImageType::kColor2D;
vk::Extent3D extent = {1, 1, 1};
vk::Format pixel_format = vk::Format::eUndefined;
uint32_t guest_format = 0;
Prospero::ImageType type = Prospero::ImageType::kColor2D;
vk::Extent3D extent = {1, 1, 1};
ImageSubresources resources;
uint32_t pitch = 0;
uint32_t bytes_per_block = 0;
@@ -352,8 +351,7 @@ inline bool ImageInfo::IsDepth() const noexcept {
}
const auto transfer_bytes = DepthAspectTransferBytes(info.pixel_format);
return transfer_bytes == info.bytes_per_block ||
(info.bytes_per_block == sizeof(uint16_t) &&
transfer_bytes == sizeof(uint32_t));
(info.bytes_per_block == sizeof(uint16_t) && transfer_bytes == sizeof(uint32_t));
}
[[nodiscard]] inline VideoOutCompression
@@ -470,18 +468,13 @@ IsSupportedDisplayRenderTargetTileMode(uint32_t tile_mode) noexcept {
vk::ClearColorValue& clear) {
vk::ClearColorValue next {};
const auto unorm8 = [](uint32_t value) { return static_cast<float>(value & 0xffu) / 255.0f; };
const auto srgb8 = [](uint32_t value) {
const auto srgb8 = [](uint32_t value) {
const auto encoded = static_cast<float>(value & 0xffu) / 255.0f;
return encoded <= 0.04045f ? encoded / 12.92f
: std::pow((encoded + 0.055f) / 1.055f, 2.4f);
return encoded <= 0.04045f ? encoded / 12.92f : std::pow((encoded + 0.055f) / 1.055f, 2.4f);
};
switch (format) {
case vk::Format::eR32Uint:
next.uint32[0] = packed;
break;
case vk::Format::eR32Sint:
next.int32[0] = static_cast<int32_t>(packed);
break;
case vk::Format::eR32Uint: next.uint32[0] = packed; break;
case vk::Format::eR32Sint: next.int32[0] = static_cast<int32_t>(packed); break;
case vk::Format::eR8G8B8A8Srgb:
next.float32[0] = srgb8(packed);
next.float32[1] = srgb8(packed >> 8u);
@@ -70,15 +70,14 @@ namespace {
}
case vk::ImageType::e3D:
switch (info.type) {
case vk::ImageViewType::e3D:
return info.base_layer == 0 && info.layer_count == 1;
case vk::ImageViewType::e3D: return info.base_layer == 0 && info.layer_count == 1;
case vk::ImageViewType::e2D:
return static_cast<bool>(
image.flags & vk::ImageCreateFlagBits::e2DArrayCompatible) &&
return static_cast<bool>(image.flags &
vk::ImageCreateFlagBits::e2DArrayCompatible) &&
info.level_count == 1 && info.layer_count == 1;
case vk::ImageViewType::e2DArray:
return static_cast<bool>(
image.flags & vk::ImageCreateFlagBits::e2DArrayCompatible) &&
return static_cast<bool>(image.flags &
vk::ImageCreateFlagBits::e2DArrayCompatible) &&
info.level_count == 1;
default: return false;
}
@@ -325,11 +324,10 @@ bool FormatsCompatible(vk::Format base, vk::Format view) noexcept {
} // namespace ImageViewOps
vk::ImageView Image::FindView(const ImageViewInfo& view_info) {
const auto& image = backing;
const auto& image = backing;
auto normalized = view_info;
const bool is_storage =
static_cast<bool>(normalized.usage & vk::ImageUsageFlagBits::eStorage);
normalized.aspect = FullAspectMask(image.format);
const bool is_storage = static_cast<bool>(normalized.usage & vk::ImageUsageFlagBits::eStorage);
normalized.aspect = FullAspectMask(image.format);
if (normalized.aspect & vk::ImageAspectFlagBits::eDepth &&
IsDepthViewFormat(normalized.format)) {
normalized.format = image.format;
@@ -340,28 +338,26 @@ vk::ImageView Image::FindView(const ImageViewInfo& view_info) {
normalized.format = image.format;
normalized.aspect = vk::ImageAspectFlagBits::eStencil;
}
normalized.usage =
is_storage ? vk::ImageUsageFlagBits::eStorage : vk::ImageUsageFlags {};
normalized.usage = is_storage ? vk::ImageUsageFlagBits::eStorage : vk::ImageUsageFlags {};
const bool format_compatible = normalized.format != vk::Format::eUndefined &&
IsCompatibleViewFormat(image.format, normalized.format);
const bool slice_view = image.image_type == vk::ImageType::e3D &&
(normalized.type == vk::ImageViewType::e2D ||
normalized.type == vk::ImageViewType::e2DArray);
const bool slice_view =
image.image_type == vk::ImageType::e3D && (normalized.type == vk::ImageViewType::e2D ||
normalized.type == vk::ImageViewType::e2DArray);
const bool levels_valid = normalized.level_count != 0 &&
normalized.base_level < image.mip_levels &&
normalized.level_count <= image.mip_levels - normalized.base_level;
const auto view_layers = slice_view && levels_valid
? std::max(image.extent.depth >> normalized.base_level, 1u)
: image.layers;
const bool ranges_valid = levels_valid &&
normalized.layer_count != 0 && normalized.base_layer < view_layers &&
const auto view_layers = slice_view && levels_valid
? std::max(image.extent.depth >> normalized.base_level, 1u)
: image.layers;
const bool ranges_valid = levels_valid && normalized.layer_count != 0 &&
normalized.base_layer < view_layers &&
normalized.layer_count <= view_layers - normalized.base_layer;
const bool mapping_valid =
IsComponentSwizzle(normalized.mapping.r) && IsComponentSwizzle(normalized.mapping.g) &&
IsComponentSwizzle(normalized.mapping.b) && IsComponentSwizzle(normalized.mapping.a);
if (image.image == nullptr || !format_compatible || !ranges_valid || !mapping_valid ||
!IsValidViewType(image, normalized) ||
!IsValidAspect(image, normalized.aspect)) {
!IsValidViewType(image, normalized) || !IsValidAspect(image, normalized.aspect)) {
EXIT("invalid image view: image_format=%d view_format=%d type=%d aspect=0x%x "
"mip=%u+%u layer=%u+%u usage=0x%x image_levels=%u image_layers=%u\n",
static_cast<int>(image.format), static_cast<int>(normalized.format),
@@ -88,7 +88,9 @@ SelectSampledDepthView(vk::Format image_format, vk::Format view_format, uint32_t
IsSupportedSampledDepthResource(const ShaderRecompiler::IR::ImageResource& resource) noexcept {
return resource.kind == ShaderRecompiler::IR::ResourceKind::Image &&
(resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2D ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DArray) &&
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DArray ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa ||
resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray) &&
resource.mip_mode == ShaderRecompiler::IR::ImageMipMode::None && resource.read &&
!resource.written && !resource.atomic;
}
@@ -397,10 +397,10 @@ TextureUploadLayout TextureCalcUploadLayout(uint32_t fmt, uint64_t width, uint64
return layout;
}
std::vector<vk::BufferImageCopy>
TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels, bool array_texture,
bool volume_texture) {
std::vector<vk::BufferImageCopy> TextureBuildImageCopies(const TextureUploadLayout& layout,
uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels,
bool array_texture, bool volume_texture) {
uint32_t mip_width = width;
uint32_t mip_height = height;
uint32_t mip_pitch = volume_texture && static_cast<Prospero::TileMode>(layout.tile) !=
@@ -416,14 +416,13 @@ TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint3
const auto mip_depth = GetTextureLevelDepth(depth, i, volume_texture);
for (uint32_t z = 0; z < mip_depth; z++) {
const auto slice_offset = z * layout.slice_stride;
const auto slice_offset = z * layout.slice_stride;
vk::BufferImageCopy region {};
region.bufferOffset =
layout.level_sizes[i].offset + slice_offset;
region.imageSubresource = {vk::ImageAspectFlagBits::eColor, i,
array_texture ? z : 0, 1};
region.imageOffset.z = volume_texture ? static_cast<int>(z) : 0;
region.imageExtent = {mip_width, mip_height, 1};
region.bufferOffset = layout.level_sizes[i].offset + slice_offset;
region.imageSubresource = {vk::ImageAspectFlagBits::eColor, i, array_texture ? z : 0,
1};
region.imageOffset.z = volume_texture ? static_cast<int>(z) : 0;
region.imageExtent = {mip_width, mip_height, 1};
const bool linear =
static_cast<Prospero::TileMode>(layout.tile) == Prospero::TileMode::kLinear;
if (linear) {
@@ -433,9 +432,8 @@ TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint3
const auto align = [](uint32_t value, uint32_t block) {
return ((value + block - 1u) / block) * block;
};
const auto pitch = align(mip_pitch, layout.texel_block);
region.bufferRowLength =
pitch > align(mip_width, layout.texel_block) ? pitch : 0;
const auto pitch = align(mip_pitch, layout.texel_block);
region.bufferRowLength = pitch > align(mip_width, layout.texel_block) ? pitch : 0;
}
regions.push_back(region);
}
@@ -480,8 +478,7 @@ static bool SetGpuTileSize(uint64_t offset, uint64_t length, uint64_t capacity,
return true;
}
bool TextureBuildGpuTileInfos(uint64_t size,
const std::vector<vk::BufferImageCopy>& regions,
bool TextureBuildGpuTileInfos(uint64_t size, const std::vector<vk::BufferImageCopy>& regions,
const TextureUploadLayout& layout, uint32_t fmt, uint32_t depth,
uint64_t levels, std::vector<GpuTileInfo>& out_infos) {
if (size == 0 || levels == 0 || levels > 16 || depth == 0 ||
@@ -522,13 +519,12 @@ bool TextureBuildGpuTileInfos(uint64_t size,
for (uint32_t z = 0; z < mip_depth; z += block.block_depth) {
const uint32_t copy_depth = std::min(block.block_depth, mip_depth - z);
const auto& region = regions[region_base + z];
const auto pitch = region.bufferRowLength != 0
? region.bufferRowLength
: region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
GpuTileInfo info {};
const auto pitch =
region.bufferRowLength != 0 ? region.bufferRowLength : region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
GpuTileInfo info {};
info.family = block.family;
info.bytes_per_element = block.bytes_per_element;
info.linear_offset = region.bufferOffset;
@@ -544,20 +540,17 @@ bool TextureBuildGpuTileInfos(uint64_t size,
return false;
}
info.linear_slice_stride = linear_stride;
info.width = std::max(
(region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max(
(logical_height + element.tall - 1u) / element.tall, 1u);
info.depth = copy_depth;
info.surface_z = block.block_depth == 1
? static_cast<uint32_t>(region.imageOffset.z)
: 0;
info.pitch =
std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.tail_x = tail ? volume.tail_x[level] : 0;
info.tail_y = tail ? volume.tail_y[level] : 0;
info.tail = tail;
info.tiled_width = volume.level_widths[level];
info.width =
std::max((region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max((logical_height + element.tall - 1u) / element.tall, 1u);
info.depth = copy_depth;
info.surface_z =
block.block_depth == 1 ? static_cast<uint32_t>(region.imageOffset.z) : 0;
info.pitch = std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.tail_x = tail ? volume.tail_x[level] : 0;
info.tail_y = tail ? volume.tail_y[level] : 0;
info.tail = tail;
info.tiled_width = volume.level_widths[level];
info.tiled_height = volume.level_heights[level];
infos.push_back(info);
}
@@ -581,12 +574,11 @@ bool TextureBuildGpuTileInfos(uint64_t size,
const auto level_depth = GetTextureLevelDepth(depth, level, layout.volume_texture);
for (uint32_t z = 0; z < level_depth; z++) {
const auto& region = regions[region_index++];
const auto pitch = region.bufferRowLength != 0
? region.bufferRowLength
: region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
const auto pitch =
region.bufferRowLength != 0 ? region.bufferRowLength : region.imageExtent.width;
const auto logical_height = region.bufferImageHeight != 0
? region.bufferImageHeight
: region.imageExtent.height;
GpuTileInfo info {};
info.family = block.family;
info.bytes_per_element = block.bytes_per_element;
@@ -597,16 +589,14 @@ bool TextureBuildGpuTileInfos(uint64_t size,
info.tiled_size)) {
return false;
}
info.width = std::max(
(region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max(
(logical_height + element.tall - 1u) / element.tall, 1u);
info.width =
std::max((region.imageExtent.width + element.wide - 1u) / element.wide, 1u);
info.height = std::max((logical_height + element.tall - 1u) / element.tall, 1u);
info.surface_z = base_family == TileBlockFamily::RenderTarget64KB ||
base_family == TileBlockFamily::Depth64KB
? region.imageSubresource.baseArrayLayer
: 0;
info.pitch =
std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.pitch = std::max((pitch + element.wide - 1u) / element.wide, 1u);
info.tail = tail;
info.tail_x = tail ? level_size.x : 0;
info.tail_y = tail ? level_size.y : 0;
@@ -32,20 +32,19 @@ struct TextureUploadLayout {
TilePaddedSize padded_sizes[16] = {};
};
vk::ComponentMapping TextureGetComponentMapping(uint32_t swizzle);
vk::ComponentMapping TextureGetComponentMapping(uint32_t swizzle);
vk::Format TextureGetFormat(uint32_t fmt);
RenderTargetFormatInfo TextureGetRenderTargetFormat(uint32_t layout, uint32_t type, uint32_t order);
TextureUploadLayout TextureCalcUploadLayout(uint32_t fmt, uint64_t width, uint64_t height,
uint64_t levels, uint32_t depth, uint64_t pitch,
uint64_t tile, uint64_t upload_size,
bool allow_depth_tile, bool volume_texture,
const char* owner);
std::vector<vk::BufferImageCopy>
TextureBuildImageCopies(const TextureUploadLayout& layout, uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels, bool array_texture,
bool volume_texture);
bool TextureBuildGpuTileInfos(uint64_t size,
const std::vector<vk::BufferImageCopy>& regions,
TextureUploadLayout TextureCalcUploadLayout(uint32_t fmt, uint64_t width, uint64_t height,
uint64_t levels, uint32_t depth, uint64_t pitch,
uint64_t tile, uint64_t upload_size,
bool allow_depth_tile, bool volume_texture,
const char* owner);
std::vector<vk::BufferImageCopy> TextureBuildImageCopies(const TextureUploadLayout& layout,
uint32_t width, uint32_t height,
uint32_t depth, uint64_t levels,
bool array_texture, bool volume_texture);
bool TextureBuildGpuTileInfos(uint64_t size, const std::vector<vk::BufferImageCopy>& regions,
const TextureUploadLayout& layout, uint32_t fmt, uint32_t depth,
uint64_t levels, std::vector<GpuTileInfo>& infos);
@@ -14,9 +14,9 @@
#include "gpu_tiler_shaders/gpu_tiler_standard64_spv.h"
#include "gpu_tiler_shaders/gpu_tiler_swap_bgra16_spv.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/cache/streamBuffer.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
#include "graphics/host_gpu/renderer/image/image.h"
#include "graphics/host_gpu/renderer/cache/streamBuffer.h"
#include <algorithm>
#include <array>
@@ -26,8 +26,8 @@ MasterSemaphore::~MasterSemaphore() {
}
void MasterSemaphore::Refresh() {
uint64_t counter = 0;
const auto result = m_graphics.device.getSemaphoreCounterValue(m_semaphore, &counter);
uint64_t counter = 0;
const auto result = m_graphics.device.getSemaphoreCounterValue(m_semaphore, &counter);
EXIT_NOT_IMPLEMENTED(result != vk::Result::eSuccess);
auto known = m_gpu_tick.load(std::memory_order_acquire);
@@ -22,7 +22,7 @@ public:
[[nodiscard]] uint64_t KnownGpuTick() const noexcept {
return m_gpu_tick.load(std::memory_order_acquire);
}
[[nodiscard]] bool IsFree(uint64_t tick) const noexcept { return KnownGpuTick() >= tick; }
[[nodiscard]] bool IsFree(uint64_t tick) const noexcept { return KnownGpuTick() >= tick; }
[[nodiscard]] uint64_t NextTick() noexcept {
return m_current_tick.fetch_add(1, std::memory_order_release);
}
@@ -26,11 +26,15 @@ bool IsSampledImage(BindingKind kind) {
case BindingKind::Sampled1DArray:
case BindingKind::Sampled2D:
case BindingKind::Sampled2DArray:
case BindingKind::Sampled2DMsaa:
case BindingKind::Sampled2DMsaaArray:
case BindingKind::Sampled3D:
case BindingKind::SampledUint1D:
case BindingKind::SampledUint1DArray:
case BindingKind::SampledUint2D:
case BindingKind::SampledUint2DArray:
case BindingKind::SampledUint2DMsaa:
case BindingKind::SampledUint2DMsaaArray:
case BindingKind::SampledUint3D: return true;
default: return false;
}
@@ -95,7 +95,7 @@ private:
};
static vk::DescriptorImageInfo MakeImageInfo(const TextureBinding& texture);
void CreatePool();
void CreatePool();
VulkanDescriptorSet* Allocate(Stage stage, const ShaderRecompiler::IR::Program& program);
vk::DescriptorSetLayout
GetDescriptorSetLayoutInternal(Stage stage, const ShaderRecompiler::IR::Program& program);
@@ -73,6 +73,11 @@ static Prospero::ImageType TextureBaseType(Prospero::ImageType type) {
}
}
static bool IsMultisampledTexture(Prospero::ImageType type) {
return type == Prospero::ImageType::kColor2DMsaa ||
type == Prospero::ImageType::kColor2DMsaaArray;
}
static BufferView NativeStorageBuffer(RenderContext& context, CommandBuffer& command_buffer,
const ShaderBufferResource& descriptor,
const ShaderRecompiler::IR::BufferResource& resource,
@@ -159,6 +164,8 @@ static bool IsSupportedSampledColorResource(const ShaderRecompiler::IR::ImageRes
case ShaderRecompiler::Decoder::ImageDimension::Dim1DArray:
case ShaderRecompiler::Decoder::ImageDimension::Dim2D:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DArray:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray:
supported_dimension = true;
break;
default: break;
@@ -195,6 +202,22 @@ TargetTextureViewInfo ResolveTargetTextureView(const ShaderRecompiler::IR::Image
? TargetTextureViewInfo {vk::ImageViewType::e2DArray, base_layer,
image_layers - base_layer}
: TargetTextureViewInfo {};
case Prospero::ImageType::kColor2DMsaa:
return resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa &&
base_layer == 0 && image_layers == 1
? TargetTextureViewInfo {vk::ImageViewType::e2D, 0, 1}
: TargetTextureViewInfo {};
case Prospero::ImageType::kColor2DMsaaArray:
if (resource.dimension == ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa &&
base_layer == 0 && image_layers == 1) {
return {vk::ImageViewType::e2D, 0, 1};
}
return resource.dimension ==
ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray &&
base_layer < image_layers
? TargetTextureViewInfo {vk::ImageViewType::e2DArray, base_layer,
image_layers - base_layer}
: TargetTextureViewInfo {};
default: return {};
}
}
@@ -209,41 +232,64 @@ bool IsSupportedSampledVideoOutView(const ShaderRecompiler::IR::ImageResource& r
}
bool IsSupportedDepthTargetDescriptor(const ShaderTextureResource& descriptor, const Image& image) {
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto pitch = TileGetTexturePitch(descriptor.Format(), width, 1, descriptor.TileMode());
const auto type = static_cast<Prospero::ImageType>(descriptor.Type());
const bool supported_single_layer =
image.info.resources.layers == 1 && descriptor.Depth() == 0 &&
descriptor.BaseArray5() == 0 &&
(type == Prospero::ImageType::kColor2D || type == Prospero::ImageType::kColor2DArray);
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto type = static_cast<Prospero::ImageType>(descriptor.Type());
const bool multisampled = IsMultisampledTexture(type);
const auto samples = multisampled ? 1u << descriptor.LastLevel() : 1u;
const auto pitch =
multisampled ? TileGetDepthPitch(width, image.info.bytes_per_block, descriptor.LastLevel())
: TileGetTexturePitch(descriptor.Format(), width, 1, descriptor.TileMode());
const bool supported_2d = type == Prospero::ImageType::kColor2D &&
image.info.resources.layers == 1 && descriptor.Depth() == 0 &&
descriptor.BaseArray5() == 0;
const bool supported_array = type == Prospero::ImageType::kColor2DArray &&
descriptor.BaseArray5() <= descriptor.Depth() &&
descriptor.Depth() < image.info.resources.layers;
const bool supported_cube =
type == Prospero::ImageType::kCube && width == height && image.info.resources.layers >= 6 &&
image.info.resources.layers % 6u == 0 &&
static_cast<uint32_t>(descriptor.Depth()) + 1u == image.info.resources.layers &&
descriptor.BaseArray5() == 0;
const bool supported_msaa_2d = type == Prospero::ImageType::kColor2DMsaa &&
image.info.resources.layers == 1 && descriptor.Depth() == 0 &&
descriptor.BaseArray5() == 0;
const bool supported_msaa_array = type == Prospero::ImageType::kColor2DMsaaArray &&
descriptor.BaseArray5() <= descriptor.Depth() &&
descriptor.Depth() < image.info.resources.layers;
const bool levels_ok =
multisampled
? descriptor.BaseLevel() == 0 && descriptor.LastLevel() >= 1 &&
descriptor.LastLevel() <= 3 && descriptor.MaxMip() == descriptor.LastLevel() &&
image.info.resources.levels == 1 && image.info.samples == samples
: descriptor.BaseLevel() == 0 && descriptor.LastLevel() == 0 &&
descriptor.MaxMip() == 0 && image.info.samples == 1;
return image.info.IsDepth() && width == image.info.extent.width &&
height == image.info.extent.height && (supported_single_layer || supported_cube) &&
descriptor.BaseLevel() == 0 && descriptor.LastLevel() == 0 && descriptor.MaxMip() == 0 &&
descriptor.MinLod() == 0 && descriptor.BaseArray5() == 0 &&
height == image.info.extent.height &&
(supported_2d || supported_array || supported_cube || supported_msaa_2d ||
supported_msaa_array) &&
levels_ok && descriptor.MinLod() == 0 &&
descriptor.TileMode() == Prospero::GpuEnumValue(Prospero::TileMode::kDepth) &&
descriptor.BCSwizzle() == 0 && !descriptor.MsaaDepth() && pitch >= width &&
pitch == image.info.pitch;
descriptor.BCSwizzle() == 0 && descriptor.MsaaDepth() == multisampled &&
pitch >= width && pitch == image.info.pitch;
}
bool IsSupportedDepthTextureEncoding(const ShaderTextureResource& descriptor, const Image& image) {
constexpr uint32_t field1_reserved_mask = 0x200fff00u;
constexpr uint32_t field2_reserved_mask = 0xf0003000u;
constexpr uint32_t field3_common = 0x01800000u;
constexpr uint32_t field5_expected = 0x00700000u;
const uint32_t field3_expected =
(descriptor.Type() << 28u) | field3_common | descriptor.DstSelXYZW();
const uint32_t field4_expected = descriptor.Depth() | (descriptor.BaseArray5() << 16u);
const bool common = (descriptor.fields[1] & field1_reserved_mask) == 0 &&
(descriptor.fields[2] & field2_reserved_mask) == 0 &&
descriptor.fields[3] == field3_expected &&
descriptor.fields[4] == field4_expected &&
descriptor.fields[5] == field5_expected;
const uint32_t field3_expected = descriptor.DstSelXYZW() |
(static_cast<uint32_t>(descriptor.BaseLevel()) << 12u) |
(static_cast<uint32_t>(descriptor.LastLevel()) << 16u) |
(static_cast<uint32_t>(descriptor.TileMode()) << 20u) |
(static_cast<uint32_t>(descriptor.Type()) << 28u);
const uint32_t field4_expected = descriptor.Depth() | (descriptor.BaseArray5() << 16u);
const uint32_t field5_expected =
0x00700000u | (static_cast<uint32_t>(descriptor.MaxMip()) << 4u);
const bool common = (descriptor.fields[1] & field1_reserved_mask) == 0 &&
(descriptor.fields[2] & field2_reserved_mask) == 0 &&
descriptor.fields[3] == field3_expected &&
descriptor.fields[4] == field4_expected &&
descriptor.fields[5] == field5_expected;
if (!common || (descriptor.fields[6] == 0 && descriptor.fields[7] != 0)) {
return false;
}
@@ -251,8 +297,9 @@ bool IsSupportedDepthTextureEncoding(const ShaderTextureResource& descriptor, co
return true;
}
constexpr uint32_t htile_control = 0x00280000u;
const auto metadata_addr = descriptor.MetaAddr() << 8u;
return (descriptor.fields[6] & 0x00ffffffu) == htile_control && metadata_addr != 0 &&
const uint32_t expected_control = htile_control | (descriptor.MsaaDepth() ? (1u << 10u) : 0u);
const auto metadata_addr = descriptor.MetaAddr() << 8u;
return (descriptor.fields[6] & 0x00ffffffu) == expected_control && metadata_addr != 0 &&
descriptor.TileMode() == Prospero::GpuEnumValue(Prospero::TileMode::kDepth) &&
image.info.tile_mode == Prospero::GpuEnumValue(Prospero::TileMode::kDepth) &&
image.info.metadata.kind == ImageMetadataKind::Htile &&
@@ -518,6 +565,7 @@ static ImageViewInfo TextureViewInfo(const ShaderRecompiler::IR::ImageResource&
view.layer_count = 1;
break;
case ShaderRecompiler::Decoder::ImageDimension::Dim2DArray:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaaArray:
view.type = vk::ImageViewType::e2DArray;
view.base_layer = descriptor.BaseArray5();
if (view.base_layer >= image_layers) {
@@ -526,6 +574,7 @@ static ImageViewInfo TextureViewInfo(const ShaderRecompiler::IR::ImageResource&
view.layer_count = image_layers - view.base_layer;
break;
case ShaderRecompiler::Decoder::ImageDimension::Dim2D:
case ShaderRecompiler::Decoder::ImageDimension::Dim2DMsaa:
view.type = vk::ImageViewType::e2D;
view.base_layer = descriptor.BaseArray5();
if (view.base_layer >= image_layers) {
@@ -556,22 +605,23 @@ RenderExecutor::ResolveTexture(const ShaderRecompiler::IR::ImageResource& reso
return {id, nullptr, std::move(desc)};
}
const auto address = descriptor.Base40();
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto base_level = descriptor.BaseLevel();
const auto last_level = descriptor.LastLevel();
const auto type = TextureType(descriptor);
const bool multisampled =
type == Prospero::ImageType::kColor2DMsaa || type == Prospero::ImageType::kColor2DMsaaArray;
const auto levels = multisampled ? 1u : static_cast<uint32_t>(descriptor.MaxMip()) + 1u;
const auto tile = descriptor.TileMode();
const bool msaa_tile = tile == Prospero::GpuEnumValue(Prospero::TileMode::kRenderTarget);
const auto address = descriptor.Base40();
const auto width = static_cast<uint32_t>(descriptor.Width5()) + 1u;
const auto height = static_cast<uint32_t>(descriptor.Height5()) + 1u;
const auto base_level = descriptor.BaseLevel();
const auto last_level = descriptor.LastLevel();
const auto type = TextureType(descriptor);
const bool multisampled = IsMultisampledTexture(type);
const auto levels = multisampled ? 1u : static_cast<uint32_t>(descriptor.MaxMip()) + 1u;
const auto tile = descriptor.TileMode();
const bool msaa_tile =
tile == Prospero::GpuEnumValue(descriptor.MsaaDepth() ? Prospero::TileMode::kDepth
: Prospero::TileMode::kRenderTarget);
const bool msaa_array = type == Prospero::ImageType::kColor2DMsaaArray;
if ((!multisampled && (base_level > last_level || last_level >= levels)) ||
(multisampled &&
(base_level != 0 || last_level == 0 || last_level > 3 ||
descriptor.MaxMip() != last_level || !msaa_tile || descriptor.MsaaDepth() ||
descriptor.MaxMip() != last_level || !msaa_tile ||
(!msaa_array && (descriptor.Depth() != 0 || descriptor.BaseArray5() != 0))))) {
EXIT("unsupported texture mip view: base=%u last=%u levels=%u\n", base_level, last_level,
levels);
@@ -36,7 +36,7 @@ ResolveTargetTextureView(const ShaderRecompiler::IR::ImageResource& resource,
[[nodiscard]] bool IsSupportedDepthTargetDescriptor(const ShaderTextureResource& descriptor,
const Image& image);
[[nodiscard]] bool IsSupportedDepthTextureEncoding(const ShaderTextureResource& descriptor,
const Image& image);
const Image& image);
[[nodiscard]] bool
IsSupportedSampledVideoOutView(const ShaderRecompiler::IR::ImageResource& resource,
const ShaderTextureResource& descriptor, const Image& image);
@@ -88,12 +88,12 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
PipelineStaticParameters static_params {};
GraphicsPipeline p {};
p.ps_shader_id = ps_id;
p.vs_shader_id = vs_id;
p.ps_shader_id = ps_id;
p.vs_shader_id = vs_id;
static_params.color_count = color_count;
PipelineRenderingState rendering {};
rendering.color_count = color_count;
rendering.color_count = color_count;
uint32_t attachment_samples = 0;
for (uint32_t i = 0; i < color_count; i++) {
EXIT_IF(!colors[i].image_id || colors[i].format == vk::Format::eUndefined);
@@ -116,8 +116,8 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
if (attachment_samples == 0) {
attachment_samples = depth.samples;
} else if (attachment_samples != depth.samples) {
EXIT("mixed color/depth sample counts are unsupported: %u and %u\n",
attachment_samples, depth.samples);
EXIT("mixed color/depth sample counts are unsupported: %u and %u\n", attachment_samples,
depth.samples);
}
}
EXIT_IF(attachment_samples == 0 ||
@@ -179,10 +179,10 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
NormalizeStaticParamsForDynamicState(static_params);
GraphicsPipelineKey key {};
key.rendering = rendering;
key.vs_shader_id = p.vs_shader_id;
key.ps_shader_id = p.ps_shader_id;
key.static_params = static_params;
key.rendering = rendering;
key.vs_shader_id = p.vs_shader_id;
key.ps_shader_id = p.ps_shader_id;
key.static_params = static_params;
if (auto iter = m_graphics_pipelines.find(key); iter != m_graphics_pipelines.end()) {
return *iter->second;
@@ -203,9 +203,8 @@ PipelineCache::GraphicsPipeline& PipelineCache::CreateGraphicsPipeline(
LogPipelineTrace("CreatePipelineInternal begin", vs_id.hash0, vs_id.crc32, ps_id.hash0,
ps_id.crc32);
CreatePipelineInternal(m_graphics, m_descriptor_cache, *cached, rendering, vs_input_info,
vs_spirv, ps_input_info,
ps_spirv, static_params, vs_id.hash0, vs_id.crc32, ps_id.hash0,
ps_id.crc32, ps_active);
vs_spirv, ps_input_info, ps_spirv, static_params, vs_id.hash0,
vs_id.crc32, ps_id.hash0, ps_id.crc32, ps_active);
LogPipelineTrace("CreatePipelineInternal done", vs_id.hash0, vs_id.crc32, ps_id.hash0,
ps_id.crc32);
@@ -88,9 +88,9 @@ static_assert(sizeof(PipelineStaticParameters) ==
struct PipelineRenderingState {
std::array<vk::Format, RENDER_COLOR_ATTACHMENTS_MAX> color_formats {};
vk::Format depth_format = vk::Format::eUndefined;
vk::Format stencil_format = vk::Format::eUndefined;
uint32_t color_count = 0;
vk::Format depth_format = vk::Format::eUndefined;
vk::Format stencil_format = vk::Format::eUndefined;
uint32_t color_count = 0;
bool operator==(const PipelineRenderingState&) const = default;
};
@@ -118,11 +118,12 @@ public:
ShaderId cs_shader_id;
};
GraphicsPipeline& CreateGraphicsPipeline(
RenderColorInfo* colors, uint32_t color_count, RenderDepthInfo& depth,
ShaderVertexInputInfo& vs_input_info, RenderCommandBuffer& command,
ShaderPixelInputInfo* ps_input_info, vk::PrimitiveTopology topology, bool ps_active,
std::span<const uint32_t> vs_spirv, std::span<const uint32_t> ps_spirv);
GraphicsPipeline&
CreateGraphicsPipeline(RenderColorInfo* colors, uint32_t color_count, RenderDepthInfo& depth,
ShaderVertexInputInfo& vs_input_info, RenderCommandBuffer& command,
ShaderPixelInputInfo* ps_input_info, vk::PrimitiveTopology topology,
bool ps_active, std::span<const uint32_t> vs_spirv,
std::span<const uint32_t> ps_spirv);
ComputePipeline& CreateComputePipeline(ShaderComputeInputInfo& input_info,
const HW::ComputeShaderInfo& cs_regs,
std::span<const uint32_t> cs_spirv);
@@ -199,7 +200,7 @@ private:
}
};
GraphicContext& m_graphics;
GraphicContext& m_graphics;
DescriptorCache& m_descriptor_cache;
std::unordered_map<GraphicsPipelineKey, std::unique_ptr<GraphicsPipeline>,
GraphicsPipelineKeyHash>
@@ -211,16 +212,13 @@ private:
void LogPipelineTrace(const char* phase, uint32_t vs_hash0, uint32_t vs_crc32, uint32_t ps_hash0,
uint32_t ps_crc32);
void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline,
const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info,
std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info,
std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0,
uint32_t vs_crc32, uint32_t ps_hash0, uint32_t ps_crc32,
bool ps_active);
void CreatePipelineInternal(
GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline, const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info, std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info, std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0, uint32_t vs_crc32,
uint32_t ps_hash0, uint32_t ps_crc32, bool ps_active);
void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::ComputePipeline& pipeline,
const ShaderComputeInputInfo& input_info,
@@ -8,10 +8,10 @@
#include "graphics/host_gpu/renderer/debug.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "graphics/host_gpu/renderer/pipeline/shaderSubgroup.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/renderer/renderTarget.h"
#include "graphics/host_gpu/renderer/pipeline/shaderSubgroup.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include "graphics/shader/recompiler/ir/ShaderIR.h"
#include "graphics/shader/shader.h"
@@ -385,9 +385,8 @@ static vk::BlendOp GetBlendOp(uint32_t op) {
return vk::BlendOp::eAdd;
}
static void CreateLayout(DescriptorCache& descriptor_cache,
std::span<vk::DescriptorSetLayout> set_layouts,
uint32_t& set_layouts_num,
static void CreateLayout(DescriptorCache& descriptor_cache,
std::span<vk::DescriptorSetLayout> set_layouts, uint32_t& set_layouts_num,
std::span<vk::PushConstantRange> push_constant_info,
uint32_t& push_constant_info_num,
const ShaderRecompiler::IR::Program& program,
@@ -412,12 +411,11 @@ static void CreateLayout(DescriptorCache& descriptor_cache,
}
}
static void ConfigureSubgroupSize(const GraphicContext& graphics,
vk::ShaderStageFlagBits vk_stage,
static void ConfigureSubgroupSize(const GraphicContext& graphics, vk::ShaderStageFlagBits vk_stage,
const ShaderRecompiler::IR::Program& program,
vk::PipelineShaderStageRequiredSubgroupSizeCreateInfo& required,
vk::PipelineShaderStageCreateInfo& stage) {
const auto config =
const auto config =
ConfigureShaderSubgroup(ShaderSubgroupCapabilities {graphics}, vk_stage, program);
switch (config.mode) {
case ShaderSubgroupMode::Natural: return;
@@ -456,16 +454,13 @@ static void ConfigureSubgroupSize(const GraphicContext&
}
// NOLINTNEXTLINE(readability-function-cognitive-complexity)
void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline,
const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info,
std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info,
std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0,
uint32_t vs_crc32, uint32_t ps_hash0, uint32_t ps_crc32,
bool ps_active) {
void CreatePipelineInternal(
GraphicContext& graphics, DescriptorCache& descriptor_cache,
PipelineCache::GraphicsPipeline& pipeline, const PipelineRenderingState& rendering,
const ShaderVertexInputInfo& vs_input_info, std::span<const uint32_t> vs_shader,
const ShaderPixelInputInfo* ps_input_info, std::span<const uint32_t> ps_shader,
const PipelineStaticParameters& static_params, uint32_t vs_hash0, uint32_t vs_crc32,
uint32_t ps_hash0, uint32_t ps_crc32, bool ps_active) {
EXIT_IF(ps_active && ps_input_info == nullptr);
vk::ShaderModule vert_shader_module = nullptr;
@@ -511,8 +506,7 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
vert_shader_stage_info.pName = "main";
vert_shader_stage_info.pSpecializationInfo = nullptr;
EXIT_IF(!vs_input_info.stage);
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eVertex,
*vs_input_info.stage.program,
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eVertex, *vs_input_info.stage.program,
vert_subgroup_size, vert_shader_stage_info);
vk::PipelineShaderStageCreateInfo frag_shader_stage_info {};
@@ -527,8 +521,8 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
if (ps_active) {
EXIT_IF(!ps_input_info->stage);
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eFragment,
*ps_input_info->stage.program,
frag_subgroup_size, frag_shader_stage_info);
*ps_input_info->stage.program, frag_subgroup_size,
frag_shader_stage_info);
}
vk::PipelineShaderStageCreateInfo shader_stages[] = {vert_shader_stage_info,
@@ -728,13 +722,13 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
clip_ext.depthClipEnable = static_params.depth_clip_enable ? VK_TRUE : VK_FALSE;
vk::PipelineRasterizationStateCreateInfo rasterizer {};
rasterizer.sType = vk::StructureType::ePipelineRasterizationStateCreateInfo;
rasterizer.sType = vk::StructureType::ePipelineRasterizationStateCreateInfo;
// MoltenVK lacks VK_EXT_depth_clip_enable; omit the depth-clip struct on macOS and accept
// Vulkan's default depth clipping (enabled) instead of the PS5's clamp behavior.
#if defined(__APPLE__)
rasterizer.pNext = nullptr;
rasterizer.pNext = nullptr;
#else
rasterizer.pNext = &clip_ext;
rasterizer.pNext = &clip_ext;
#endif
rasterizer.flags = {};
rasterizer.depthClampEnable = VK_FALSE;
@@ -812,13 +806,13 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
color_write.pColorWriteEnables = color_write_enable;
vk::PipelineColorBlendStateCreateInfo color_blending {};
color_blending.sType = vk::StructureType::ePipelineColorBlendStateCreateInfo;
color_blending.sType = vk::StructureType::ePipelineColorBlendStateCreateInfo;
// MoltenVK lacks VK_EXT_color_write_enable; drop the dynamic color-write struct on macOS
// and rely on each attachment's static colorWriteMask (all channels enabled by default).
#if defined(__APPLE__)
color_blending.pNext = nullptr;
color_blending.pNext = nullptr;
#else
color_blending.pNext = &color_write;
color_blending.pNext = &color_write;
#endif
color_blending.flags = {};
color_blending.logicOpEnable = VK_FALSE;
@@ -838,15 +832,13 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
EXIT_IF(!vs_input_info.stage);
CreateLayout(descriptor_cache, set_layouts, set_layouts_num, push_constant_info,
push_constant_info_num,
*vs_input_info.stage.program, vk::ShaderStageFlagBits::eVertex,
DescriptorCache::Stage::Vertex);
push_constant_info_num, *vs_input_info.stage.program,
vk::ShaderStageFlagBits::eVertex, DescriptorCache::Stage::Vertex);
if (ps_active) {
EXIT_IF(!ps_input_info->stage);
CreateLayout(descriptor_cache, set_layouts, set_layouts_num, push_constant_info,
push_constant_info_num,
*ps_input_info->stage.program, vk::ShaderStageFlagBits::eFragment,
DescriptorCache::Stage::Pixel);
push_constant_info_num, *ps_input_info->stage.program,
vk::ShaderStageFlagBits::eFragment, DescriptorCache::Stage::Pixel);
}
vk::PipelineLayoutCreateInfo pipeline_layout_info {};
@@ -923,32 +915,32 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
dynamic_state.dynamicStateCount = dynamic_states_count;
dynamic_state.pDynamicStates = dynamic_states;
vk::GraphicsPipelineCreateInfo pipeline_info {};
vk::GraphicsPipelineCreateInfo pipeline_info {};
vk::PipelineRenderingCreateInfo rendering_info {};
rendering_info.sType = vk::StructureType::ePipelineRenderingCreateInfo;
rendering_info.colorAttachmentCount = rendering.color_count;
rendering_info.pColorAttachmentFormats = rendering.color_formats.data();
rendering_info.depthAttachmentFormat = rendering.depth_format;
rendering_info.stencilAttachmentFormat = rendering.stencil_format;
pipeline_info.sType = vk::StructureType::eGraphicsPipelineCreateInfo;
pipeline_info.pNext = &rendering_info;
pipeline_info.flags = {};
pipeline_info.stageCount = shader_stage_count;
pipeline_info.pStages = shader_stages;
pipeline_info.pVertexInputState = &vertex_input_info;
pipeline_info.pInputAssemblyState = &input_assembly;
pipeline_info.pTessellationState = nullptr;
pipeline_info.pViewportState = &viewport_state;
pipeline_info.pRasterizationState = &rasterizer;
pipeline_info.pMultisampleState = &multisampling;
pipeline_info.pDepthStencilState = (static_params.with_depth ? &depth_stencil_info : nullptr);
pipeline_info.pColorBlendState = &color_blending;
pipeline_info.pDynamicState = &dynamic_state;
pipeline_info.layout = pipeline.pipeline_layout;
pipeline_info.renderPass = nullptr;
pipeline_info.subpass = 0;
pipeline_info.basePipelineHandle = nullptr;
pipeline_info.basePipelineIndex = -1;
pipeline_info.sType = vk::StructureType::eGraphicsPipelineCreateInfo;
pipeline_info.pNext = &rendering_info;
pipeline_info.flags = {};
pipeline_info.stageCount = shader_stage_count;
pipeline_info.pStages = shader_stages;
pipeline_info.pVertexInputState = &vertex_input_info;
pipeline_info.pInputAssemblyState = &input_assembly;
pipeline_info.pTessellationState = nullptr;
pipeline_info.pViewportState = &viewport_state;
pipeline_info.pRasterizationState = &rasterizer;
pipeline_info.pMultisampleState = &multisampling;
pipeline_info.pDepthStencilState = (static_params.with_depth ? &depth_stencil_info : nullptr);
pipeline_info.pColorBlendState = &color_blending;
pipeline_info.pDynamicState = &dynamic_state;
pipeline_info.layout = pipeline.pipeline_layout;
pipeline_info.renderPass = nullptr;
pipeline_info.subpass = 0;
pipeline_info.basePipelineHandle = nullptr;
pipeline_info.basePipelineIndex = -1;
EXIT_IF(pipeline.pipeline != nullptr);
@@ -1012,8 +1004,7 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
comp_shader_stage_info.pName = "main";
comp_shader_stage_info.pSpecializationInfo = nullptr;
EXIT_IF(!input_info.stage);
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eCompute,
*input_info.stage.program,
ConfigureSubgroupSize(graphics, vk::ShaderStageFlagBits::eCompute, *input_info.stage.program,
comp_subgroup_size, comp_shader_stage_info);
vk::DescriptorSetLayout set_layouts[1] = {};
@@ -1024,9 +1015,8 @@ void CreatePipelineInternal(GraphicContext& graphics, DescriptorCache& descripto
EXIT_IF(!input_info.stage);
CreateLayout(descriptor_cache, set_layouts, set_layouts_num, push_constant_info,
push_constant_info_num,
*input_info.stage.program, vk::ShaderStageFlagBits::eCompute,
DescriptorCache::Stage::Compute);
push_constant_info_num, *input_info.stage.program,
vk::ShaderStageFlagBits::eCompute, DescriptorCache::Stage::Compute);
vk::PipelineLayoutCreateInfo pipeline_layout_info {};
pipeline_layout_info.sType = vk::StructureType::ePipelineLayoutCreateInfo;
@@ -10,14 +10,14 @@
#include "graphics/guest_gpu/graphicsRun.h"
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/host_gpu/graphicContext.h"
#include "graphics/host_gpu/renderer/image/imageInfo.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/pipeline/descriptors.h"
#include "graphics/host_gpu/renderer/image/imageInfo.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/renderer/pipeline/shaderResourceBarrier.h"
#include "graphics/host_gpu/renderer/pipeline/shaderSubgroup.h"
#include "graphics/host_gpu/renderer/render.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include "graphics/shader/recompiler/ir/ResourceMaterialization.h"
#include "graphics/shader/recompiler/ir/ShaderIR.h"
@@ -14,8 +14,7 @@ namespace Libs::Graphics {
RenderContext::RenderContext(GraphicContext& graphics)
: m_graphics(graphics), m_render_executor(*this), m_command_scheduler(*this, graphics),
m_descriptor_cache(graphics), m_pipeline_cache(graphics, m_descriptor_cache),
m_sampler_cache(graphics),
m_gpu_resources(graphics, m_command_scheduler) {
m_sampler_cache(graphics), m_gpu_resources(graphics, m_command_scheduler) {
EXIT_NOT_IMPLEMENTED(!Common::Thread::IsMainThread());
}
@@ -27,7 +26,7 @@ RenderContext::~RenderContext() {
void RenderContext::InitializeGpu(VideoOut::VideoOutDriver* video_out) {
EXIT_IF(m_gpu != nullptr);
m_video_out = video_out;
m_gpu = std::make_unique<Gpu>(*this);
m_gpu = std::make_unique<Gpu>(*this);
m_gpu_resources.SetGpu(m_gpu.get());
}
@@ -99,8 +98,7 @@ void RenderContext::TriggerEopEvent(uint32_t context_id) {
registration.eq, static_cast<uintptr_t>(registration.id),
LibKernel::EventQueue::KERNEL_EVFILT_GRAPHICS,
reinterpret_cast<void*>(static_cast<uintptr_t>(context_id)));
if (result == LibKernel::KERNEL_ERROR_EBADF ||
result == LibKernel::KERNEL_ERROR_ENOENT) {
if (result == LibKernel::KERNEL_ERROR_EBADF || result == LibKernel::KERNEL_ERROR_ENOENT) {
DeleteEopEq(registration.eq, registration.id);
continue;
}
+17 -17
View File
@@ -6,12 +6,12 @@
#include "common/common.h"
#include "common/threads.h"
#include "graphics/host_gpu/renderer/cache/bufferCache.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/cache/gpuResourceManager.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "graphics/host_gpu/renderer/cache/samplerCache.h"
#include "graphics/host_gpu/renderer/cache/textureCache.h"
#include "graphics/host_gpu/renderer/commandScheduler.h"
#include "graphics/host_gpu/renderer/pipeline/descriptorCache.h"
#include "graphics/host_gpu/renderer/pipeline/pipelineCache.h"
#include "kernel/eventQueue.h"
#include <memory>
@@ -32,10 +32,10 @@ public:
~RenderContext();
KYTY_CLASS_NO_COPY(RenderContext);
[[nodiscard]] GraphicContext& GetGraphics() const noexcept { return m_graphics; }
void InitializeGpu(VideoOut::VideoOutDriver* video_out);
void ShutdownGpu();
[[nodiscard]] Gpu& GetGpu() const;
[[nodiscard]] GraphicContext& GetGraphics() const noexcept { return m_graphics; }
void InitializeGpu(VideoOut::VideoOutDriver* video_out);
void ShutdownGpu();
[[nodiscard]] Gpu& GetGpu() const;
[[nodiscard]] VideoOut::VideoOutDriver& GetVideoOut() const;
Common::Mutex& GetMutex() { return m_mutex; }
@@ -56,18 +56,18 @@ private:
struct EopEqRegistration {
LibKernel::EventQueue::KernelEqueue eq = LibKernel::EventQueue::KERNEL_EQUEUE_INVALID;
LibKernel::EventQueue::KernelEqueueRef queue;
int id = 0;
int id = 0;
};
GraphicContext& m_graphics;
Common::Mutex m_mutex;
RenderExecutor m_render_executor;
CommandScheduler m_command_scheduler;
DescriptorCache m_descriptor_cache;
PipelineCache m_pipeline_cache;
SamplerCache m_sampler_cache;
GpuResourceManager m_gpu_resources;
std::unique_ptr<Gpu> m_gpu;
GraphicContext& m_graphics;
Common::Mutex m_mutex;
RenderExecutor m_render_executor;
CommandScheduler m_command_scheduler;
DescriptorCache m_descriptor_cache;
PipelineCache m_pipeline_cache;
SamplerCache m_sampler_cache;
GpuResourceManager m_gpu_resources;
std::unique_ptr<Gpu> m_gpu;
VideoOut::VideoOutDriver* m_video_out = nullptr;
Common::Mutex m_eop_mutex;
@@ -12,13 +12,13 @@ namespace Libs::Graphics {
static constexpr uint32_t RENDER_COLOR_ATTACHMENTS_MAX = 8;
struct RenderAttachment {
vk::ImageView image_view = nullptr;
vk::ImageLayout image_layout = vk::ImageLayout::eUndefined;
std::array<uint32_t, 4> clear_value = {};
vk::ImageView image_view = nullptr;
vk::ImageLayout image_layout = vk::ImageLayout::eUndefined;
std::array<uint32_t, 4> clear_value = {};
bool is_clear = false;
bool has_depth = false;
bool depth_clear = false;
bool has_stencil = false;
bool has_depth = false;
bool depth_clear = false;
bool has_stencil = false;
bool stencil_clear = false;
bool operator==(const RenderAttachment&) const = default;
+3 -3
View File
@@ -251,9 +251,9 @@ uint64_t PrepareVideoOutFlip(CommandBuffer& buffer, int handle, int index, int f
int64_t flip_arg) {
for (;;) {
uint64_t request_id = 0;
auto& video_out = buffer.GetContext().GetVideoOut();
const auto result = video_out.SubmitFlipFromGpu(
buffer, handle, index, flip_mode, flip_arg, request_id);
auto& video_out = buffer.GetContext().GetVideoOut();
const auto result =
video_out.SubmitFlipFromGpu(buffer, handle, index, flip_mode, flip_arg, request_id);
if (result == OK) {
EXIT_IF(request_id == 0);
return request_id;
+6 -7
View File
@@ -122,9 +122,9 @@ uint64_t GraphicContext::GetDeviceMemoryUsage() const {
physical_device_properties.deviceType == vk::PhysicalDeviceType::eDiscreteGpu;
uint64_t usage = 0;
for (uint32_t heap = 0; heap < physical_device_memory_properties.memoryHeapCount; heap++) {
const bool device_local = static_cast<bool>(
physical_device_memory_properties.memoryHeaps[heap].flags &
vk::MemoryHeapFlagBits::eDeviceLocal);
const bool device_local =
static_cast<bool>(physical_device_memory_properties.memoryHeaps[heap].flags &
vk::MemoryHeapFlagBits::eDeviceLocal);
if (!discrete || device_local) {
usage += budgets[heap].usage;
}
@@ -144,7 +144,7 @@ uint64_t GraphicContext::GetTotalMemoryBudget() const {
uint64_t local = 0;
uint64_t usage = 0;
for (uint32_t heap = 0; heap < physical_device_memory_properties.memoryHeapCount; heap++) {
const auto& properties = physical_device_memory_properties.memoryHeaps[heap];
const auto& properties = physical_device_memory_properties.memoryHeaps[heap];
const bool device_local =
static_cast<bool>(properties.flags & vk::MemoryHeapFlagBits::eDeviceLocal);
if (device_local) {
@@ -159,9 +159,8 @@ uint64_t GraphicContext::GetTotalMemoryBudget() const {
return budget - std::min<uint64_t>(budget / 8, 1024ull * 1024 * 1024);
}
constexpr uint64_t system_reserve = 8ull * 1024 * 1024 * 1024;
const auto available = budget > usage ? budget - usage : uint64_t {0};
return std::max(local,
available > system_reserve ? available - system_reserve : uint64_t {0});
const auto available = budget > usage ? budget - usage : uint64_t {0};
return std::max(local, available > system_reserve ? available - system_reserve : uint64_t {0});
}
void GraphicContext::CreateBuffer(uint64_t size, VulkanBuffer& buffer) {
+4
View File
@@ -55,6 +55,10 @@ constexpr FormatMapping kFormatMappings[] = {
{Prospero::BufferFormat::k32_32_32_32UInt, vk::Format::eR32G32B32A32Uint},
{Prospero::BufferFormat::k32_32_32_32SInt, vk::Format::eR32G32B32A32Sint},
{Prospero::BufferFormat::k32_32_32_32Float, vk::Format::eR32G32B32A32Sfloat},
// Narrow-channel sRGB formats are optional in Vulkan. Keep a same-width fallback until
// sampler-aware sRGB emulation is available.
{Prospero::BufferFormat::k8Srgb, vk::Format::eR8Unorm},
{Prospero::BufferFormat::k8_8Srgb, vk::Format::eR8G8Unorm},
{Prospero::BufferFormat::k8_8_8_8Srgb, vk::Format::eR8G8B8A8Srgb},
{Prospero::BufferFormat::k9_9_9_5Float, vk::Format::eE5B9G9R9UfloatPack32},
{Prospero::BufferFormat::k5_6_5UNorm, vk::Format::eB5G6R5UnormPack16},
+7 -7
View File
@@ -20,14 +20,14 @@ public:
~Presenter();
KYTY_CLASS_NO_COPY(Presenter);
[[nodiscard]] Frame& PrepareFrame(CommandBuffer& command, const ImageInfo& info);
[[nodiscard]] Frame& PrepareBlankFrame(uint32_t width, uint32_t height, bool opaque,
CommandBuffer* producer = nullptr);
[[nodiscard]] Frame* PrepareLastFrame();
[[nodiscard]] bool IsGuestPaused() const noexcept;
[[nodiscard]] Frame& PrepareFrame(CommandBuffer& command, const ImageInfo& info);
[[nodiscard]] Frame& PrepareBlankFrame(uint32_t width, uint32_t height, bool opaque,
CommandBuffer* producer = nullptr);
[[nodiscard]] Frame* PrepareLastFrame();
[[nodiscard]] bool IsGuestPaused() const noexcept;
[[nodiscard]] RenderContext& Renderer() const noexcept;
void Present(Frame& frame, bool reuse = false);
void Discard(Frame& frame);
void Present(Frame& frame, bool reuse = false);
void Discard(Frame& frame);
private:
struct Impl;
+37 -41
View File
@@ -69,8 +69,8 @@ enum class FlipRequestSource { Cpu, GpuEop };
struct VideoOutEventState;
struct VideoOutEventRegistration {
EventQueue::KernelEqueue handle = EventQueue::KERNEL_EQUEUE_INVALID;
std::shared_ptr<VideoOutEventState> state;
EventQueue::KernelEqueue handle = EventQueue::KERNEL_EQUEUE_INVALID;
std::shared_ptr<VideoOutEventState> state;
uint64_t generation = 0;
VideoOutEventKind kind = VideoOutEventKind::Flip;
};
@@ -170,13 +170,13 @@ struct BufferAttributeGroup {
struct VideoOutConfig {
Common::Mutex mutex;
Common::CondVar vblank_cond;
std::shared_ptr<VideoOutEventState> events = std::make_shared<VideoOutEventState>();
uint32_t width = 0;
uint32_t height = 0;
uint64_t generation = 0;
bool opened = false;
bool closing = false;
int flip_rate = 0;
std::shared_ptr<VideoOutEventState> events = std::make_shared<VideoOutEventState>();
uint32_t width = 0;
uint32_t height = 0;
uint64_t generation = 0;
bool opened = false;
bool closing = false;
int flip_rate = 0;
uint64_t output_mode = VIDEO_OUT_OUTPUT_MODE_DEFAULT;
float gamma = 1.0f;
VideoOutFlipStatus flip_status;
@@ -250,8 +250,8 @@ public:
VideoOutConfig* Get(int handle, uint64_t& generation);
bool IsOpened(int handle);
void Init(uint32_t width, uint32_t height);
FlipQueue& GetFlipQueue() { return m_flip_queue; }
void Init(uint32_t width, uint32_t height);
FlipQueue& GetFlipQueue() { return m_flip_queue; }
Graphics::RenderContext& Renderer() const noexcept { return m_renderer; }
void VblankBegin();
@@ -259,12 +259,12 @@ public:
void PresentThread(std::stop_token token);
private:
Common::Mutex m_mutex;
VideoOutConfig m_video_out_ctx[VIDEO_OUT_NUM_MAX];
Common::Mutex m_mutex;
VideoOutConfig m_video_out_ctx[VIDEO_OUT_NUM_MAX];
Graphics::RenderContext& m_renderer;
Graphics::Presenter& m_presenter;
FlipQueue m_flip_queue;
std::jthread m_present_thread;
Graphics::Presenter& m_presenter;
FlipQueue m_flip_queue;
std::jthread m_present_thread;
};
static std::unique_ptr<VideoOutDriver> g_video_out_driver;
@@ -279,7 +279,7 @@ static uintptr_t VideoOutEventId(VideoOutEventKind kind) {
}
static VideoOutEventQueues& VideoOutEventQueuesFor(VideoOutEventState& state,
VideoOutEventKind kind) {
VideoOutEventKind kind) {
switch (kind) {
case VideoOutEventKind::Flip: return state.flip;
case VideoOutEventKind::Vblank: return state.vblank;
@@ -359,9 +359,9 @@ static void TriggerVideoOutEvents(VideoOutConfig& video_out, VideoOutEventKind k
if (!registration || registration->generation != video_out.generation) {
continue;
}
const auto result = EventQueue::KernelTriggerEvent(
registration->handle, VideoOutEventId(kind), EventQueue::KERNEL_EVFILT_VIDEO_OUT,
trigger_data);
const auto result =
EventQueue::KernelTriggerEvent(registration->handle, VideoOutEventId(kind),
EventQueue::KERNEL_EVFILT_VIDEO_OUT, trigger_data);
EXIT_NOT_IMPLEMENTED(result != OK && result != LibKernel::KERNEL_ERROR_EBADF &&
result != LibKernel::KERNEL_ERROR_ENOENT);
}
@@ -372,9 +372,8 @@ static void DeleteVideoOutEvents(const VideoOutEventQueues& queues, VideoOutEven
if (!registration) {
continue;
}
const auto result =
EventQueue::KernelDeleteEvent(registration->handle, VideoOutEventId(kind),
EventQueue::KERNEL_EVFILT_VIDEO_OUT);
const auto result = EventQueue::KernelDeleteEvent(
registration->handle, VideoOutEventId(kind), EventQueue::KERNEL_EVFILT_VIDEO_OUT);
EXIT_NOT_IMPLEMENTED(result != OK && result != LibKernel::KERNEL_ERROR_EBADF &&
result != LibKernel::KERNEL_ERROR_ENOENT);
}
@@ -383,7 +382,7 @@ static void DeleteVideoOutEvents(const VideoOutEventQueues& queues, VideoOutEven
static int RegisterVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoOutEventKind kind,
void* udata) {
uint64_t generation = 0;
auto* video_out = DriverState().Get(handle, generation);
auto* video_out = DriverState().Get(handle, generation);
if (video_out == nullptr) {
return VIDEO_OUT_ERROR_INVALID_HANDLE;
}
@@ -425,27 +424,25 @@ static int RegisterVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoO
bool add_queue = false;
{
Common::LockGuard event_lock(event_state->mutex);
const auto existing = std::find_if(queues.begin(), queues.end(), [&](const auto& candidate) {
return candidate->handle == eq && candidate->generation == generation;
});
const auto existing =
std::find_if(queues.begin(), queues.end(), [&](const auto& candidate) {
return candidate->handle == eq && candidate->generation == generation;
});
if (existing != queues.end()) {
registration = *existing;
} else {
registration = std::make_shared<VideoOutEventRegistration>(
VideoOutEventRegistration {.handle = eq,
.state = event_state,
.generation = generation,
.kind = kind});
registration = std::make_shared<VideoOutEventRegistration>(VideoOutEventRegistration {
.handle = eq, .state = event_state, .generation = generation, .kind = kind});
queues.push_back(registration);
add_queue = true;
}
}
event.filter.data = registration.get();
event.filter.owner = registration;
const int result = EventQueue::KernelAddEvent(eq, event);
const int result = EventQueue::KernelAddEvent(eq, event);
if (result != OK && add_queue) {
Common::LockGuard event_lock(event_state->mutex);
const auto added = std::find(queues.begin(), queues.end(), registration);
const auto added = std::find(queues.begin(), queues.end(), registration);
if (added != queues.end()) {
queues.erase(added);
}
@@ -455,7 +452,7 @@ static int RegisterVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoO
static int DeleteVideoOutEvent(int handle, EventQueue::KernelEqueue eq, VideoOutEventKind kind) {
uint64_t generation = 0;
auto* video_out = DriverState().Get(handle, generation);
auto* video_out = DriverState().Get(handle, generation);
if (video_out == nullptr) {
return VIDEO_OUT_ERROR_INVALID_HANDLE;
}
@@ -814,8 +811,8 @@ void VideoOutDriver::Impl::PresentThread(std::stop_token token) {
m_presenter.Present(*frame, true);
}
const auto frame_end = Common::Timer::QueryPerformanceCounter();
total_wait += static_cast<int64_t>(period) -
static_cast<int64_t>(frame_end - frame_begin);
total_wait +=
static_cast<int64_t>(period) - static_cast<int64_t>(frame_end - frame_begin);
continue;
}
@@ -841,8 +838,7 @@ void VideoOutDriver::Impl::PresentThread(std::stop_token token) {
VblankEnd();
const auto frame_end = Common::Timer::QueryPerformanceCounter();
total_wait += static_cast<int64_t>(period) -
static_cast<int64_t>(frame_end - frame_begin);
total_wait += static_cast<int64_t>(period) - static_cast<int64_t>(frame_end - frame_begin);
}
}
@@ -1000,8 +996,8 @@ void FlipQueue::Prepare(uint64_t request_id, Graphics::CommandBuffer& buffer) {
}
Graphics::Presenter::Frame* frame = nullptr;
if (special) {
frame = &m_presenter.PrepareBlankFrame(width, height,
index == VIDEO_OUT_BUFFER_INDEX_BLACK, &buffer);
frame = &m_presenter.PrepareBlankFrame(width, height, index == VIDEO_OUT_BUFFER_INDEX_BLACK,
&buffer);
} else {
frame = &m_presenter.PrepareFrame(buffer, source_info);
}
+7 -7
View File
@@ -32,13 +32,13 @@ public:
~VideoOutDriver();
KYTY_CLASS_NO_COPY(VideoOutDriver);
int SubmitFlipFromGpu(Graphics::CommandBuffer& buffer, int handle, int index, int flip_mode,
int64_t flip_arg, uint64_t& request_id);
void PrepareFlip(uint64_t request_id, Graphics::CommandBuffer& buffer);
void CompleteFlip(uint64_t request_id);
void SubmitFlipPreparation(uint64_t request_id);
void WaitForSubmitSlot();
void WaitFlipDone(int handle, int index);
int SubmitFlipFromGpu(Graphics::CommandBuffer& buffer, int handle, int index, int flip_mode,
int64_t flip_arg, uint64_t& request_id);
void PrepareFlip(uint64_t request_id, Graphics::CommandBuffer& buffer);
void CompleteFlip(uint64_t request_id);
void SubmitFlipPreparation(uint64_t request_id);
void WaitForSubmitSlot();
void WaitFlipDone(int handle, int index);
[[nodiscard]] Impl& State() noexcept;
+61 -83
View File
@@ -61,7 +61,7 @@ namespace Libs::Graphics {
struct Presenter::Frame {
VulkanImage image;
std::unique_ptr<CommandBuffer> present_commands;
bool busy = false;
bool busy = false;
bool reusing_last = false;
void Configure(GraphicContext& graphics, vk::Extent2D extent, vk::Format format);
@@ -155,7 +155,7 @@ public:
EXIT("last submitted frame is not available for reuse\n");
}
m_free.erase(free);
m_last_frame = nullptr;
m_last_frame = nullptr;
frame->busy = true;
frame->reusing_last = true;
m_mutex.Unlock();
@@ -197,30 +197,27 @@ private:
}
}
WindowContext& m_window;
Common::Mutex m_mutex;
Common::CondVar m_available;
WindowContext& m_window;
Common::Mutex m_mutex;
Common::CondVar m_available;
std::vector<std::unique_ptr<Presenter::Frame>> m_frames;
std::deque<Presenter::Frame*> m_free;
Presenter::Frame* m_last_frame = nullptr;
vk::Format m_format = vk::Format::eUndefined;
vk::Format m_format = vk::Format::eUndefined;
};
void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent,
vk::Format format) {
void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent, vk::Format format) {
if (extent.width == 0 || extent.height == 0 || format == vk::Format::eUndefined) {
EXIT("unsupported prepared frame, extent=%ux%u format=%d\n", extent.width, extent.height,
static_cast<int>(format));
}
const auto features = graphics.GetFormatProperties(format).optimalTilingFeatures;
const auto required = vk::FormatFeatureFlagBits::eBlitSrc |
vk::FormatFeatureFlagBits::eSampledImageFilterLinear |
vk::FormatFeatureFlagBits::eTransferSrc |
vk::FormatFeatureFlagBits::eTransferDst;
const auto required =
vk::FormatFeatureFlagBits::eBlitSrc | vk::FormatFeatureFlagBits::eSampledImageFilterLinear |
vk::FormatFeatureFlagBits::eTransferSrc | vk::FormatFeatureFlagBits::eTransferDst;
if ((features & required) != required) {
EXIT("prepared presentation format lacks optimal blit support: format=%d features=0x%x\n",
static_cast<int>(format),
static_cast<vk::FormatFeatureFlags::MaskType>(features));
static_cast<int>(format), static_cast<vk::FormatFeatureFlags::MaskType>(features));
}
auto& dst = image;
@@ -234,11 +231,11 @@ void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent,
dst.memory = {};
}
dst.extent = {extent.width, extent.height, 1};
dst.format = format;
dst.layers = 1;
dst.mip_levels = 1;
dst.state = {};
dst.extent = {extent.width, extent.height, 1};
dst.format = format;
dst.layers = 1;
dst.mip_levels = 1;
dst.state = {};
dst.subresource_states.clear();
dst.memory.property = vk::MemoryPropertyFlagBits::eDeviceLocal;
@@ -262,13 +259,12 @@ void Presenter::Frame::Configure(GraphicContext& graphics, vk::Extent2D extent,
void Presenter::Frame::Transit(vk::CommandBuffer command, vk::ImageLayout layout,
vk::AccessFlags2 access) {
const auto stage = access == vk::AccessFlagBits2::eTransferRead ||
access == vk::AccessFlagBits2::eTransferWrite
? vk::PipelineStageFlagBits2::eTransfer
: vk::PipelineStageFlagBits2::eAllCommands;
const auto stage = access == vk::AccessFlagBits2::eTransferRead ||
access == vk::AccessFlagBits2::eTransferWrite
? vk::PipelineStageFlagBits2::eTransfer
: vk::PipelineStageFlagBits2::eAllCommands;
constexpr auto writes = vk::AccessFlagBits2::eTransferWrite |
vk::AccessFlagBits2::eShaderWrite |
vk::AccessFlagBits2::eMemoryWrite;
vk::AccessFlagBits2::eShaderWrite | vk::AccessFlagBits2::eMemoryWrite;
if (image.state.layout == layout && image.state.access_mask == access &&
!static_cast<bool>(image.state.access_mask & writes)) {
return;
@@ -299,35 +295,27 @@ void Presenter::Frame::Transit(vk::CommandBuffer command, vk::ImageLayout layout
void Presenter::Frame::CopyFrom(CommandBuffer& command_buffer, Image& source) {
command_buffer.EndRendering();
auto command = command_buffer.Handle();
source.Transit(vk::ImageLayout::eTransferSrcOptimal,
vk::AccessFlagBits2::eTransferRead, {}, command);
Transit(command, vk::ImageLayout::eTransferDstOptimal,
vk::AccessFlagBits2::eTransferWrite);
source.Transit(vk::ImageLayout::eTransferSrcOptimal, vk::AccessFlagBits2::eTransferRead, {},
command);
Transit(command, vk::ImageLayout::eTransferDstOptimal, vk::AccessFlagBits2::eTransferWrite);
vk::ImageCopy copy {};
copy.srcSubresource = {vk::ImageAspectFlagBits::eColor, 0, 0,
source.backing.layers};
copy.srcSubresource = {vk::ImageAspectFlagBits::eColor, 0, 0, source.backing.layers};
copy.dstSubresource = {vk::ImageAspectFlagBits::eColor, 0, 0, image.layers};
copy.extent = {std::min(source.backing.extent.width, image.extent.width),
std::min(source.backing.extent.height, image.extent.height), 1};
copy.extent = {std::min(source.backing.extent.width, image.extent.width),
std::min(source.backing.extent.height, image.extent.height), 1};
EXIT_IF(copy.srcSubresource.layerCount != copy.dstSubresource.layerCount);
command.copyImage(source.backing.image, vk::ImageLayout::eTransferSrcOptimal,
image.image, vk::ImageLayout::eTransferDstOptimal, copy);
Transit(command, vk::ImageLayout::eTransferSrcOptimal,
vk::AccessFlagBits2::eTransferRead);
command.copyImage(source.backing.image, vk::ImageLayout::eTransferSrcOptimal, image.image,
vk::ImageLayout::eTransferDstOptimal, copy);
Transit(command, vk::ImageLayout::eTransferSrcOptimal, vk::AccessFlagBits2::eTransferRead);
}
void Presenter::Frame::Clear(CommandBuffer& command_buffer,
const vk::ClearColorValue& color) {
void Presenter::Frame::Clear(CommandBuffer& command_buffer, const vk::ClearColorValue& color) {
command_buffer.EndRendering();
auto command = command_buffer.Handle();
Transit(command, vk::ImageLayout::eTransferDstOptimal,
vk::AccessFlagBits2::eTransferWrite);
const vk::ImageSubresourceRange range {
vk::ImageAspectFlagBits::eColor, 0, 1, 0, 1};
command.clearColorImage(image.image, vk::ImageLayout::eTransferDstOptimal, &color, 1,
&range);
Transit(command, vk::ImageLayout::eTransferSrcOptimal,
vk::AccessFlagBits2::eTransferRead);
Transit(command, vk::ImageLayout::eTransferDstOptimal, vk::AccessFlagBits2::eTransferWrite);
const vk::ImageSubresourceRange range {vk::ImageAspectFlagBits::eColor, 0, 1, 0, 1};
command.clearColorImage(image.image, vk::ImageLayout::eTransferDstOptimal, &color, 1, &range);
Transit(command, vk::ImageLayout::eTransferSrcOptimal, vk::AccessFlagBits2::eTransferRead);
}
class Swapchain final {
@@ -338,8 +326,8 @@ public:
~Swapchain();
KYTY_CLASS_NO_COPY(Swapchain);
void Create();
void Recreate(bool surface_lost = false);
void Create();
void Recreate(bool surface_lost = false);
[[nodiscard]] Status AcquireNextImage();
void RecordPresentCommands(CommandBuffer& command, VulkanImage& source);
void Submit(CommandBuffer& command);
@@ -395,17 +383,17 @@ struct Presenter::Impl {
desc.view_info.usage = vk::ImageUsageFlagBits::eTransferSrc;
desc.type = TextureCache::BindingType::VideoOut;
auto& cache = renderer.GetTextureCache();
auto& image = cache.GetImage(cache.FindImage(desc));
auto& cache = renderer.GetTextureCache();
auto& image = cache.GetImage(cache.FindImage(desc));
image.usage.video_out = true;
return image;
}
RenderContext& renderer;
WindowContext& window;
Swapchain swapchain;
RenderContext& renderer;
WindowContext& window;
Swapchain swapchain;
CommandScheduler present_scheduler;
FramePool frames;
FramePool frames;
};
void Swapchain::Create() {
@@ -441,25 +429,20 @@ void Swapchain::Create() {
? vk::CompositeAlphaFlagBitsKHR::eOpaque
: vk::CompositeAlphaFlagBitsKHR::eInherit;
vk::SurfaceFormatKHR format {vk::Format::eR8G8B8A8Unorm,
vk::ColorSpaceKHR::eSrgbNonlinear};
if (surface.formats.size() != 1 ||
surface.formats.front().format != vk::Format::eUndefined) {
vk::SurfaceFormatKHR format {vk::Format::eR8G8B8A8Unorm, vk::ColorSpaceKHR::eSrgbNonlinear};
if (surface.formats.size() != 1 || surface.formats.front().format != vk::Format::eUndefined) {
const auto it = std::find_if(surface.formats.begin(), surface.formats.end(),
[](const vk::SurfaceFormatKHR& candidate) {
return candidate.format ==
vk::Format::eB8G8R8A8Unorm ||
candidate.format ==
vk::Format::eR8G8B8A8Unorm;
return candidate.format == vk::Format::eB8G8R8A8Unorm ||
candidate.format == vk::Format::eR8G8B8A8Unorm;
});
if (it == surface.formats.end()) {
EXIT("no supported UNORM swapchain format\n");
}
format = *it;
}
m_format = format.format;
const auto swapchain_features =
graphics.GetFormatProperties(m_format).optimalTilingFeatures;
m_format = format.format;
const auto swapchain_features = graphics.GetFormatProperties(m_format).optimalTilingFeatures;
if (!static_cast<bool>(swapchain_features & vk::FormatFeatureFlagBits::eBlitDst)) {
EXIT("swapchain format cannot be a blit destination: format=%d\n",
static_cast<int>(m_format));
@@ -503,9 +486,8 @@ void Swapchain::Create() {
view.subresourceRange.baseMipLevel = 0;
view.subresourceRange.layerCount = 1;
view.subresourceRange.levelCount = 1;
RequireVulkanSuccess(
graphics.device.createImageView(&view, nullptr, &m_image_views[i]),
"vkCreateImageView");
RequireVulkanSuccess(graphics.device.createImageView(&view, nullptr, &m_image_views[i]),
"vkCreateImageView");
EXIT_IF(m_image_views[i] == nullptr);
}
@@ -600,7 +582,7 @@ void Swapchain::Recreate(bool surface_lost) {
Swapchain::Status Swapchain::AcquireNextImage() {
EXIT_IF(m_handle == nullptr || m_frame_index >= m_image_acquired.size());
m_image_index = static_cast<uint32_t>(-1);
m_image_index = static_cast<uint32_t>(-1);
const auto result = m_window.graphic_ctx.device.acquireNextImageKHR(
m_handle, std::numeric_limits<uint64_t>::max(), m_image_acquired[m_frame_index], nullptr,
&m_image_index);
@@ -683,10 +665,9 @@ void Swapchain::RecordPresentCommands(CommandBuffer& command, VulkanImage& sourc
to_present.subresourceRange.levelCount = 1;
to_present.subresourceRange.baseArrayLayer = 0;
to_present.subresourceRange.layerCount = 1;
vk_command.pipelineBarrier(vk::PipelineStageFlagBits::eAllCommands,
vk::PipelineStageFlagBits::eAllCommands,
vk::DependencyFlagBits::eByRegion, 0,
nullptr, 0, nullptr, 1, &to_present);
vk_command.pipelineBarrier(
vk::PipelineStageFlagBits::eAllCommands, vk::PipelineStageFlagBits::eAllCommands,
vk::DependencyFlagBits::eByRegion, 0, nullptr, 0, nullptr, 1, &to_present);
command.End();
}
@@ -700,7 +681,7 @@ void Swapchain::Submit(CommandBuffer& command) {
Swapchain::Status Swapchain::Present() {
EXIT_IF(m_image_index >= m_render_complete.size());
const auto ready = m_render_complete[m_image_index];
const auto ready = m_render_complete[m_image_index];
vk::PresentInfoKHR present {};
present.sType = vk::StructureType::ePresentInfoKHR;
present.swapchainCount = 1;
@@ -738,7 +719,7 @@ Presenter::~Presenter() = default;
Presenter::Frame& Presenter::PrepareFrame(CommandBuffer& buffer, const ImageInfo& info) {
KYTY_PROFILER_FUNCTION();
EXIT_IF(buffer.IsInvalid());
auto* frame = m_impl->frames.Acquire();
auto* frame = m_impl->frames.Acquire();
Common::LockGuard render_lock(m_impl->renderer.GetMutex());
auto& image = m_impl->ResolveSurface(info);
if (image.backing.format == vk::Format::eUndefined) {
@@ -752,14 +733,13 @@ Presenter::Frame& Presenter::PrepareFrame(CommandBuffer& buffer, const ImageInfo
default: break;
}
frame->Configure(m_impl->window.graphic_ctx,
{image.backing.extent.width, image.backing.extent.height},
frame_format);
{image.backing.extent.width, image.backing.extent.height}, frame_format);
frame->CopyFrom(buffer, image);
return *frame;
}
Presenter::Frame& Presenter::PrepareBlankFrame(uint32_t width, uint32_t height, bool opaque,
CommandBuffer* producer) {
CommandBuffer* producer) {
KYTY_PROFILER_FUNCTION();
auto format = m_impl->frames.GetFormat();
auto* frame = m_impl->frames.Acquire();
@@ -772,8 +752,7 @@ Presenter::Frame& Presenter::PrepareBlankFrame(uint32_t width, uint32_t height,
frame->Clear(*producer, clear);
} else {
if (frame->present_commands == nullptr) {
frame->present_commands =
std::make_unique<CommandBuffer>(m_impl->present_scheduler);
frame->present_commands = std::make_unique<CommandBuffer>(m_impl->present_scheduler);
}
auto& command = *frame->present_commands;
command.WaitForFenceAndReset();
@@ -830,8 +809,7 @@ void Presenter::Present(Frame& frame, bool reuse) {
continue;
}
if (frame.present_commands == nullptr) {
frame.present_commands =
std::make_unique<CommandBuffer>(m_impl->present_scheduler);
frame.present_commands = std::make_unique<CommandBuffer>(m_impl->present_scheduler);
}
{
Common::LockGuard render_lock(m_impl->renderer.GetMutex());
@@ -32,11 +32,11 @@
#include "graphics/host_gpu/vma.h"
#include "graphics/host_gpu/vulkanCommon.h"
#include "graphics/presentation/presenter.h"
#include "kernel/memory.h"
#include "graphics/presentation/renderDoc.h"
#include "graphics/presentation/videoOut.h"
#include "graphics/presentation/window.h"
#include "graphics/presentation/window/windowInternal.h"
#include "kernel/memory.h"
#include "libs/controller.h"
#include "loader/systemContent.h"
@@ -475,9 +475,9 @@ static void VulkanInitSubgroupSizeControl(vk::PhysicalDevice physical_device,
}
static vk::Device VulkanCreateDevice(vk::PhysicalDevice physical_device, const VulkanExtensions& r,
uint32_t queue_family,
uint32_t queue_family,
const std::vector<const char*>& device_extensions,
GraphicContext& graphics) {
GraphicContext& graphics) {
EXIT_IF(physical_device == nullptr);
EXIT_IF(queue_family == static_cast<uint32_t>(-1));
@@ -551,19 +551,19 @@ static vk::Device VulkanCreateDevice(vk::PhysicalDevice physical_device, const V
features12.timelineSemaphore = VK_TRUE;
vk::PhysicalDeviceFeatures device_features {};
device_features.fragmentStoresAndAtomics = VK_TRUE;
device_features.samplerAnisotropy = VK_TRUE;
device_features.robustBufferAccess = VK_TRUE;
device_features.fragmentStoresAndAtomics = VK_TRUE;
device_features.samplerAnisotropy = VK_TRUE;
device_features.robustBufferAccess = VK_TRUE;
#if !defined(__APPLE__)
device_features.depthBounds = VK_TRUE; // unsupported by MoltenVK
device_features.depthBounds = VK_TRUE; // unsupported by MoltenVK
#endif
device_features.shaderStorageImageWriteWithoutFormat = VK_TRUE;
device_features.shaderStorageImageReadWithoutFormat = VK_TRUE;
device_features.shaderImageGatherExtended = VK_TRUE;
device_features.independentBlend = VK_TRUE;
device_features.tessellationShader = VK_TRUE;
device_features.sampleRateShading = VK_TRUE;
graphics.sample_rate_shading_enabled = true;
device_features.sampleRateShading = VK_TRUE;
graphics.sample_rate_shading_enabled = true;
device_features.vertexPipelineStoresAndAtomics =
supported_features2.features.vertexPipelineStoresAndAtomics;
@@ -909,10 +909,9 @@ void WindowContext::CreateVulkan() {
}
surface = native_surface;
std::vector<const char*> device_extensions = {VK_KHR_SWAPCHAIN_EXTENSION_NAME,
VK_EXT_DEPTH_CLIP_CONTROL_EXTENSION_NAME,
VK_KHR_PUSH_DESCRIPTOR_EXTENSION_NAME,
"VK_KHR_maintenance1"};
std::vector<const char*> device_extensions = {
VK_KHR_SWAPCHAIN_EXTENSION_NAME, VK_EXT_DEPTH_CLIP_CONTROL_EXTENSION_NAME,
VK_KHR_PUSH_DESCRIPTOR_EXTENSION_NAME, "VK_KHR_maintenance1"};
#if defined(__APPLE__)
// MoltenVK lacks VK_EXT_depth_clip_enable and VK_EXT_color_write_enable; the renderer
@@ -932,8 +931,8 @@ void WindowContext::CreateVulkan() {
uint32_t queue_family = static_cast<uint32_t>(-1);
VulkanFindPhysicalDevice(graphic_ctx.instance, surface, device_extensions,
surface_capabilities, graphic_ctx.physical_device, queue_family);
VulkanFindPhysicalDevice(graphic_ctx.instance, surface, device_extensions, surface_capabilities,
graphic_ctx.physical_device, queue_family);
if (graphic_ctx.physical_device == nullptr) {
EXIT("Could not find suitable device");
@@ -949,9 +948,8 @@ void WindowContext::CreateVulkan() {
auto available_extensions = EnumerateVulkan<vk::ExtensionProperties>(
"vkEnumerateDeviceExtensionProperties",
[&](uint32_t* count, vk::ExtensionProperties* values) {
return graphic_ctx.physical_device.enumerateDeviceExtensionProperties(nullptr,
count,
values);
return graphic_ctx.physical_device.enumerateDeviceExtensionProperties(
nullptr, count, values);
});
if (HasExtension(available_extensions, VK_EXT_MEMORY_BUDGET_EXTENSION_NAME)) {
@@ -985,7 +983,7 @@ void WindowContext::CreateVulkan() {
render_context = std::make_unique<RenderContext>(graphic_ctx);
LibKernel::Memory::InstallGpuResources(&render_context->GetGpuResources());
presenter = std::make_unique<Presenter>(*this);
presenter = std::make_unique<Presenter>(*this);
RenderDocSetActiveWindow(graphic_ctx.instance, window);
}
+23 -25
View File
@@ -1,7 +1,5 @@
#include "graphics/presentation/window.h"
#include <cstdlib>
#include "SDL.h"
#include "SDL_error.h"
#include "SDL_events.h"
@@ -40,6 +38,7 @@
#include <algorithm>
#include <cstdio>
#include <cstdlib>
#include <cstring>
#include <memory>
#include <string>
@@ -59,7 +58,7 @@
namespace Libs::Graphics {
constexpr int KEYBOARD_CONTROLLER_ID = -1000;
constexpr int KEYBOARD_CONTROLLER_ID = -1000;
struct EventKeyboard {
bool down;
@@ -251,9 +250,7 @@ static void GameEventKeyboard(WindowLoopState& game, const EventKeyboard& key) {
if (key.down) {
switch (key.key_code) {
case SDLK_ESCAPE: game.need_exit = true; break;
case SDLK_SPACE:
SetPause(game, !game.paused.load(std::memory_order_acquire));
break;
case SDLK_SPACE: SetPause(game, !game.paused.load(std::memory_order_acquire)); break;
case SDLK_F1:
if (!key.repeat) {
RenderDocRequestCapture();
@@ -390,7 +387,9 @@ void WindowContext::Resize(uint32_t new_width, uint32_t new_height) {
void WindowContext::ProcessWindowEvent(const SDL_WindowEvent& event) {
const auto& window_event = event;
switch (window_event.event) {
case SDL_WINDOWEVENT_SHOWN: LOGF("Window %" PRIu32 " shown\n", window_event.windowID); break;
case SDL_WINDOWEVENT_SHOWN:
LOGF("Window %" PRIu32 " shown\n", window_event.windowID);
break;
case SDL_WINDOWEVENT_HIDDEN:
LOGF("Window %" PRIu32 " hidden\n", window_event.windowID);
@@ -401,13 +400,13 @@ void WindowContext::ProcessWindowEvent(const SDL_WindowEvent& event) {
break;
case SDL_WINDOWEVENT_MOVED:
LOGF("Window %" PRIu32 " moved to %" PRId32 ",%" PRId32 "\n",
window_event.windowID, window_event.data1, window_event.data2);
LOGF("Window %" PRIu32 " moved to %" PRId32 ",%" PRId32 "\n", window_event.windowID,
window_event.data1, window_event.data2);
break;
case SDL_WINDOWEVENT_RESIZED:
LOGF("Window %" PRIu32 " resized to %" PRId32 "x%" PRId32 "\n",
window_event.windowID, window_event.data1, window_event.data2);
LOGF("Window %" PRIu32 " resized to %" PRId32 "x%" PRId32 "\n", window_event.windowID,
window_event.data1, window_event.data2);
LOGF("m: %d\n", static_cast<int>(SDL_ThreadID()));
Resize(window_event.data1, window_event.data2);
@@ -807,9 +806,8 @@ static void WindowCreate(WindowContext& context) {
window_flags |= static_cast<uint32_t>(SDL_WINDOW_BORDERLESS);
}
#endif
context.window =
SDL_CreateWindow(KYTY_SDL_WINDOW_CAPTION, KYTY_SDL_WINDOWPOS_CENTERED,
KYTY_SDL_WINDOWPOS_CENTERED, width, height, window_flags);
context.window = SDL_CreateWindow(KYTY_SDL_WINDOW_CAPTION, KYTY_SDL_WINDOWPOS_CENTERED,
KYTY_SDL_WINDOWPOS_CENTERED, width, height, window_flags);
context.window_hidden = true;
@@ -832,7 +830,7 @@ Presenter& WindowInit(uint32_t width, uint32_t height) {
WindowCreate(*window);
window->CreateVulkan();
auto& presenter = *window->presenter;
g_window = std::move(window);
g_window = std::move(window);
return presenter;
}
@@ -934,9 +932,9 @@ void WindowContext::UpdateTitle() {
Loader::SystemContentParamSfoGetString("TITLE_ID", title_id, sizeof(title_id));
static bool has_app_ver =
Loader::SystemContentParamSfoGetString("APP_VER", app_ver, sizeof(app_ver));
static uint64_t fps_start = Common::Timer::QueryPerformanceCounter();
static uint64_t frame_num = 0;
static uint64_t fps_frames = 0;
static uint64_t fps_start = Common::Timer::QueryPerformanceCounter();
static uint64_t frame_num = 0;
static uint64_t fps_frames = 0;
static double current_fps = 0.0;
const auto now = Common::Timer::QueryPerformanceCounter();
@@ -946,15 +944,15 @@ void WindowContext::UpdateTitle() {
if (now - fps_start >= frequency) {
current_fps = static_cast<double>(fps_frames) * static_cast<double>(frequency) /
static_cast<double>(now - fps_start);
fps_start = now;
fps_frames = 0;
fps_start = now;
fps_frames = 0;
}
auto fps = fmt::format("{}{}{}{}{}{}[{}] [{}], frame: {}, fps: {:f}", (has_title ? title : ""),
(has_title ? ", " : ""), (has_title_id ? title_id : ""),
(has_title_id ? ", " : ""), (has_app_ver ? app_ver : ""),
(has_app_ver ? " " : ""), device_name, processor_name,
frame_num, current_fps);
auto fps =
fmt::format("{}{}{}{}{}{}[{}] [{}], frame: {}, fps: {:f}", (has_title ? title : ""),
(has_title ? ", " : ""), (has_title_id ? title_id : ""),
(has_title_id ? ", " : ""), (has_app_ver ? app_ver : ""),
(has_app_ver ? " " : ""), device_name, processor_name, frame_num, current_fps);
#if defined(__APPLE__)
// AppKit traps on title changes off the main thread; fire-and-forget keeps present pacing.
@@ -28,8 +28,8 @@ struct SurfaceCapabilities {
};
struct WindowLoopState {
SDL_Event event {};
bool need_exit = false;
SDL_Event event {};
bool need_exit = false;
std::atomic_bool paused = false;
};
@@ -38,14 +38,13 @@ struct WindowContext {
~WindowContext();
KYTY_CLASS_NO_COPY(WindowContext);
[[nodiscard]] static vk::PhysicalDeviceVulkan13Features
RequiredVulkan13Features() noexcept;
void CreateVulkan();
void RecreateSurface();
void RefreshSurfaceCapabilities();
void UpdateIcon();
void UpdateTitle();
void Resize(uint32_t width, uint32_t height);
[[nodiscard]] static vk::PhysicalDeviceVulkan13Features RequiredVulkan13Features() noexcept;
void CreateVulkan();
void RecreateSurface();
void RefreshSurfaceCapabilities();
void UpdateIcon();
void UpdateTitle();
void Resize(uint32_t width, uint32_t height);
void ProcessWindowEvent(const SDL_WindowEvent& event);
void ProcessDisplayEvent(const SDL_DisplayEvent& event);
void ProcessEvent(double time_seconds);
@@ -59,14 +58,14 @@ struct WindowContext {
void DrainMainThreadTasks();
#endif
GraphicContext graphic_ctx;
SDL_Window* window = nullptr;
bool window_hidden = true;
vk::SurfaceKHR surface = nullptr;
SurfaceCapabilities surface_capabilities;
GraphicContext graphic_ctx;
SDL_Window* window = nullptr;
bool window_hidden = true;
vk::SurfaceKHR surface = nullptr;
SurfaceCapabilities surface_capabilities;
std::unique_ptr<RenderContext> render_context;
std::unique_ptr<Presenter> presenter;
WindowLoopState loop;
std::unique_ptr<Presenter> presenter;
WindowLoopState loop;
char device_name[VK_MAX_PHYSICAL_DEVICE_NAME_SIZE] = {0};
char processor_name[64] = {0};
@@ -76,7 +75,7 @@ struct WindowContext {
#if defined(__APPLE__)
Common::Mutex main_task_mutex;
Common::CondVar main_task_done;
std::vector<std::function<void()>> main_tasks; // guarded by main_task_mutex
std::vector<std::function<void()>> main_tasks; // guarded by main_task_mutex
uint64_t main_tasks_queued = 0; // guarded by main_task_mutex
uint64_t main_tasks_run = 0; // guarded by main_task_mutex
#endif
@@ -2,15 +2,16 @@
#include "common/assert.h"
#include "common/logging/log.h"
#include "graphics/shader/recompiler/cfg/ShaderCFG.h"
#include "graphics/shader/recompiler/decompiler/ShaderDecoder.h"
#include "graphics/shader/recompiler/emitter/SpirvEmitter.h"
#include "graphics/shader/recompiler/ir/BindingLayout.h"
#include "graphics/shader/recompiler/ir/ReadLaneElimination.h"
#include "graphics/shader/recompiler/ir/ResourceMaterialization.h"
#include "graphics/shader/recompiler/ir/ResourceTracking.h"
#include "graphics/shader/recompiler/ir/ScalarProvenance.h"
#include "graphics/shader/recompiler/cfg/ShaderCFG.h"
#include "graphics/shader/recompiler/decompiler/ShaderDecoder.h"
#include "graphics/shader/recompiler/ir/ShaderIR.h"
#include "graphics/shader/recompiler/ir/ShaderInfoCollection.h"
#include "graphics/shader/recompiler/emitter/SpirvEmitter.h"
#include "graphics/shader/recompiler/ir/SrtPatcher.h"
#include "graphics/shader/recompiler/ir/SrtWalker.h"
@@ -838,6 +839,11 @@ bool TryRecompile(std::span<const uint32_t> code, const CompileOptions& options,
if (!IR::AllocateBindings(ir, layout_options, error)) {
return false;
}
const auto read_lane_stats = IR::EliminateReadLane(ir);
if (read_lane_stats.rewritten_reads != 0) {
LOGF("%s read-lane elimination: reads=%" PRIu32 " shadow_writes=%" PRIu32 "\n",
GetDumpLabel(options), read_lane_stats.rewritten_reads, read_lane_stats.shadow_writes);
}
std::string ir_dump;
if (options.dump_ir) {
ir_dump = MakeIrDump(cfg, ir);
@@ -44,7 +44,7 @@ struct CompileResult {
};
bool TryRecompile(std::span<const uint32_t> code, const CompileOptions& options,
CompileResult& result, std::string* error);
CompileResult& result, std::string* error);
} // namespace Libs::Graphics::ShaderRecompiler
@@ -35,9 +35,9 @@ constexpr ImageDimension DecodeImageDimension(uint32_t dim) {
case 2u: return ImageDimension::Dim3D;
case 3u: return ImageDimension::Dim2DArray;
case 4u: return ImageDimension::Dim1DArray;
case 5u:
case 7u: return ImageDimension::Dim2DArray;
case 6u: return ImageDimension::Dim2D;
case 5u: return ImageDimension::Dim2DArray;
case 6u: return ImageDimension::Dim2DMsaa;
case 7u: return ImageDimension::Dim2DMsaaArray;
default: return ImageDimension::Unknown;
}
}
@@ -46,8 +46,10 @@ constexpr uint32_t ImageCoordComponents(ImageDimension dimension) {
switch (dimension) {
case ImageDimension::Dim1D: return 1u;
case ImageDimension::Dim1DArray: return 2u;
case ImageDimension::Dim2DMsaa:
case ImageDimension::Dim3D:
case ImageDimension::Dim2DArray: return 3u;
case ImageDimension::Dim2DMsaaArray: return 4u;
default: return 2u;
}
}
@@ -196,10 +196,10 @@ bool DecodeSopk(uint32_t pc, std::span<const uint32_t> code, uint32_t word_index
case Opcode::SMovkI32: return DecodeScalarDestination(sdst, pc, inst.dst, error);
case Opcode::SWaitcnt: {
const uint32_t waitcnt = word & 0xffffu;
inst.dst.kind = OperandKind::Null;
inst.src0.signed_val = static_cast<int32_t>(waitcnt);
inst.src0.value = waitcnt;
inst.src_count = 1;
inst.dst.kind = OperandKind::Null;
inst.src0.signed_val = static_cast<int32_t>(waitcnt);
inst.src0.value = waitcnt;
inst.src_count = 1;
return true;
}
case Opcode::SSetregB32:
@@ -266,10 +266,10 @@ bool DecodeSopp(uint32_t pc, std::span<const uint32_t> code, uint32_t word_index
inst.src0.value = simm;
inst.src0.signed_val = static_cast<int16_t>(simm);
inst.src_count = (inst.opcode == Opcode::SNop || inst.opcode == Opcode::SWaitcnt ||
inst.opcode == Opcode::SSleep || inst.opcode == Opcode::SSendmsg ||
inst.opcode == Opcode::STtraceData || inst.opcode == Opcode::SInstPrefetch)
? 1
: 0;
inst.opcode == Opcode::SSleep || inst.opcode == Opcode::SSendmsg ||
inst.opcode == Opcode::STtraceData || inst.opcode == Opcode::SInstPrefetch)
? 1
: 0;
inst.branch_offset = static_cast<int32_t>(static_cast<int16_t>(simm)) * 4;
inst.branch_target = pc + 4u + static_cast<uint32_t>(inst.branch_offset);
SetRawWords(inst, code, word_index, 1);
@@ -194,6 +194,8 @@ const char* ImageDimensionToString(ImageDimension dimension) {
case ImageDimension::Dim2D: return "2d";
case ImageDimension::Dim3D: return "3d";
case ImageDimension::Dim2DArray: return "2d_array";
case ImageDimension::Dim2DMsaa: return "2d_msaa";
case ImageDimension::Dim2DMsaaArray: return "2d_msaa_array";
default: return "unknown";
}
}
@@ -220,9 +222,9 @@ bool DecodeScalarSource(uint32_t code, uint32_t pc, Operand& operand, std::strin
}
if (code >= 240u && code <= 247u) {
constexpr float values[] = {0.5f, -0.5f, 1.0f, -1.0f, 2.0f, -2.0f, 4.0f, -4.0f};
operand.kind = OperandKind::FloatInlineConstant;
operand.float_val = values[code - 240u];
operand.value = FloatBits(operand.float_val);
operand.kind = OperandKind::FloatInlineConstant;
operand.float_val = values[code - 240u];
operand.value = FloatBits(operand.float_val);
return true;
}
if (code >= 256u && code <= 511u) {
@@ -285,7 +287,7 @@ bool DecodeVectorGpr(uint32_t reg, Operand& operand, std::string* error) {
SetError(error, "VGPR index is out of range");
return false;
}
operand = {};
operand = {};
operand.kind = OperandKind::Vgpr;
operand.reg = reg;
return true;
@@ -575,6 +575,8 @@ enum class ImageDimension : uint32_t {
Dim2D,
Dim3D,
Dim2DArray,
Dim2DMsaa,
Dim2DMsaaArray,
};
constexpr uint32_t MaxInstructionRawWords = 5u;
@@ -30,10 +30,16 @@ bool ImageBinding(const IR::ImageResource& image, IR::DescriptorBindingKind& kin
kind = integer ? Kind::SampledUint1DArray : Kind::Sampled1DArray;
return true;
case Dim::Dim2D: kind = integer ? Kind::SampledUint2D : Kind::Sampled2D; return true;
case Dim::Dim2DMsaa:
kind = integer ? Kind::SampledUint2DMsaa : Kind::Sampled2DMsaa;
return true;
case Dim::Dim3D: kind = integer ? Kind::SampledUint3D : Kind::Sampled3D; return true;
case Dim::Dim2DArray:
kind = integer ? Kind::SampledUint2DArray : Kind::Sampled2DArray;
return true;
case Dim::Dim2DMsaaArray:
kind = integer ? Kind::SampledUint2DMsaaArray : Kind::Sampled2DMsaaArray;
return true;
case Dim::Unknown: return false;
}
}
@@ -51,6 +57,8 @@ bool ImageBinding(const IR::ImageResource& image, IR::DescriptorBindingKind& kin
case Dim::Dim2DArray:
kind = uint_image ? Kind::StorageUint2DArray : Kind::Storage2DArray;
return true;
case Dim::Dim2DMsaa:
case Dim::Dim2DMsaaArray: return false;
case Dim::Unknown: return false;
}
return false;
@@ -12,9 +12,9 @@ namespace Libs::Graphics::ShaderRecompiler::Spirv {
bool ProgramRequiresExactSubgroupSize(const IR::Program& program);
bool EmitProgram(const IR::Program& program, const IR::ResourceSnapshot& resources,
const ShaderVertexInputInfo* vertex_input_info,
const ShaderPixelInputInfo* pixel_input_info,
const ShaderComputeInputInfo* compute_input_info, std::vector<uint32_t>& spirv,
const ShaderVertexInputInfo* vertex_input_info,
const ShaderPixelInputInfo* pixel_input_info,
const ShaderComputeInputInfo* compute_input_info, std::vector<uint32_t>& spirv,
std::string* error);
} // namespace Libs::Graphics::ShaderRecompiler::Spirv
@@ -129,7 +129,7 @@ uint32_t MaxCollectedVectorRegisterEnd(const std::vector<RegisterBinding>& regis
}
void CollectMoveRelSourceRegisters(const IR::Program& program,
std::vector<RegisterBinding>& registers) {
std::vector<RegisterBinding>& registers) {
const auto max_vector_end = MaxCollectedVectorRegisterEnd(registers);
for (const auto& block: program.blocks) {
for (const auto& inst: block.instructions) {
@@ -276,8 +276,7 @@ void CopyProgramInputsAndOutputs(EmitterState& state, const IR::Program& program
if (HasOutput(state.outputs, output.kind, output.index)) {
continue;
}
state.outputs.push_back(
{output.kind, output.index, output.location, 0, output.debug_name});
state.outputs.push_back({output.kind, output.index, output.location, 0, output.debug_name});
}
}
@@ -576,6 +575,8 @@ ImageViewKind ImageViewKindFromDimension(Decoder::ImageDimension dimension) {
case Decoder::ImageDimension::Dim1DArray: return ImageViewKind::Dim1DArray;
case Decoder::ImageDimension::Dim2DArray: return ImageViewKind::Dim2DArray;
case Decoder::ImageDimension::Dim3D: return ImageViewKind::Dim3D;
case Decoder::ImageDimension::Dim2DMsaa: return ImageViewKind::Dim2DMsaa;
case Decoder::ImageDimension::Dim2DMsaaArray: return ImageViewKind::Dim2DMsaaArray;
default: return ImageViewKind::Dim2D;
}
}
@@ -601,7 +602,9 @@ uint32_t ImageViewCoordinateComponents(ImageViewKind view) {
case ImageViewKind::Dim1DArray:
case ImageViewKind::Dim2D: return 2u;
case ImageViewKind::Dim2DArray:
case ImageViewKind::Dim2DMsaaArray:
case ImageViewKind::Dim3D: return 3u;
case ImageViewKind::Dim2DMsaa: return 2u;
default: return 0u;
}
}
@@ -611,7 +614,9 @@ uint32_t ImageViewSpatialComponents(ImageViewKind view) {
case ImageViewKind::Dim1D:
case ImageViewKind::Dim1DArray: return 1u;
case ImageViewKind::Dim2D:
case ImageViewKind::Dim2DArray: return 2u;
case ImageViewKind::Dim2DArray:
case ImageViewKind::Dim2DMsaa:
case ImageViewKind::Dim2DMsaaArray: return 2u;
case ImageViewKind::Dim3D: return 3u;
default: return 0u;
}
@@ -663,8 +668,7 @@ uint32_t LoadSampledImageDescriptor(EmitterState& state, const IR::MemoryInfo& m
uint32_t LoadSamplerDescriptor(EmitterState& state, uint32_t sampler, uint32_t use_pc) {
(void)use_pc;
const auto binding =
ResourceForDescriptor(state, IR::DescriptorBindingKind::Samplers, sampler);
const auto binding = ResourceForDescriptor(state, IR::DescriptorBindingKind::Samplers, sampler);
const auto pointer = DescriptorElementPointer(
state, state.ptr_uniform_sampler, state.sampler_variable, binding.array_index,
IR::DescriptorBindingKind::Samplers, sampler, "sampler descriptor array was not emitted");
@@ -24,7 +24,7 @@ uint32_t EmitExportVec4F32(EmitterState& state, const IR::Instruction& inst) {
const auto raw = EmitValueLoad(state, inst.src[pair_index]);
const auto unpacked = state.builder.AllocateId();
state.builder.AddFunction({OpExtInst, state.vec2_float_type, unpacked,
state.glsl_std450, GlslUnpackHalf2x16, raw});
state.glsl_std450, GlslUnpackHalf2x16, raw});
for (uint32_t lane = 0; lane < 2u; lane++) {
const auto component = pair_index * 2u + lane;
if (((inst.export_info.en >> component) & 1u) == 0) {
@@ -36,8 +36,8 @@ uint32_t EmitExportVec4F32(EmitterState& state, const IR::Instruction& inst) {
}
}
const auto vec = state.builder.AllocateId();
state.builder.AddFunction({OpCompositeConstruct, state.vec4_float_type, vec,
components[0], components[1], components[2], components[3]});
state.builder.AddFunction({OpCompositeConstruct, state.vec4_float_type, vec, components[0],
components[1], components[2], components[3]});
return vec;
}
@@ -50,7 +50,59 @@ uint32_t EmitExportVec4F32(EmitterState& state, const IR::Instruction& inst) {
return vec;
}
uint32_t ApplyMrtExportMapping(EmitterState& state, const IR::Instruction& inst, uint32_t value) {
uint32_t EmitExportComponentU32(EmitterState& state, const IR::Instruction& inst,
uint32_t component) {
const bool enabled = ((inst.export_info.en >> component) & 1u) != 0;
if (!enabled || component >= inst.src_count || component >= 4u) {
return ConstantU32(state, component == 3u ? 1u : 0u);
}
return EmitValueLoad(state, inst.src[component]);
}
uint32_t EmitExportVec4U32(EmitterState& state, const IR::Instruction& inst) {
uint32_t components[4] = {
ConstantU32(state, 0u),
ConstantU32(state, 0u),
ConstantU32(state, 0u),
ConstantU32(state, 1u),
};
if (inst.export_info.compr) {
for (uint32_t pair_index = 0; pair_index < 2u && pair_index < inst.src_count;
pair_index++) {
const auto raw = EmitValueLoad(state, inst.src[pair_index]);
for (uint32_t lane = 0; lane < 2u; lane++) {
const auto component = pair_index * 2u + lane;
if (((inst.export_info.en >> component) & 1u) == 0) {
continue;
}
components[component] = state.builder.AllocateId();
state.builder.AddFunction(
{OpBitFieldUExtract, state.uint_type, components[component], raw,
ConstantU32(state, lane * 16u), ConstantU32(state, 16u)});
}
}
} else {
for (uint32_t component = 0; component < 4u; component++) {
components[component] = EmitExportComponentU32(state, inst, component);
}
}
const auto vec = state.builder.AllocateId();
state.builder.AddFunction({OpCompositeConstruct, state.vec4_uint_type, vec, components[0],
components[1], components[2], components[3]});
return vec;
}
static bool MrtUsesUintOutput(const EmitterState& state, const IR::Instruction& inst) {
return inst.export_info.kind == IR::ExportTargetKind::Mrt &&
state.pixel_input_info != nullptr &&
inst.export_info.index < std::size(state.pixel_input_info->target_output_mode) &&
state.pixel_input_info->target_output_mode[inst.export_info.index] == 7u;
}
uint32_t ApplyMrtExportMapping(EmitterState& state, const IR::Instruction& inst, uint32_t value,
uint32_t vector_type) {
if (inst.export_info.kind != IR::ExportTargetKind::Mrt || state.pixel_input_info == nullptr ||
inst.export_info.index >= state.pixel_input_info->target_export_mapping.size()) {
return value;
@@ -62,8 +114,8 @@ uint32_t ApplyMrtExportMapping(EmitterState& state, const IR::Instruction& inst,
}
const auto mapped = state.builder.AllocateId();
state.builder.AddFunction({OpVectorShuffle, state.vec4_float_type, mapped, value, value,
mapping.Map(0), mapping.Map(1), mapping.Map(2), mapping.Map(3)});
state.builder.AddFunction({OpVectorShuffle, vector_type, mapped, value, value, mapping.Map(0),
mapping.Map(1), mapping.Map(2), mapping.Map(3)});
return mapped;
}
@@ -89,7 +141,7 @@ void EmitMrtZExport(EmitterState& state, const IR::Instruction& inst) {
const auto ptr = state.builder.AllocateId();
state.builder.AddFunction({OpBitcast, state.int_type, mask, raw});
state.builder.AddFunction({OpAccessChain, state.ptr_output_int, ptr,
state.sample_mask_variable, ConstantU32(state, 0)});
state.sample_mask_variable, ConstantU32(state, 0)});
state.builder.AddFunction({OpStore, ptr, mask});
}
}
@@ -114,11 +166,15 @@ void EmitExport(EmitterState& state, const IR::Instruction& inst) {
return;
}
const auto value = ApplyMrtExportMapping(state, inst, EmitExportVec4F32(state, inst));
const auto uint_output = MrtUsesUintOutput(state, inst);
const auto vector_type = uint_output ? state.vec4_uint_type : state.vec4_float_type;
const auto value = ApplyMrtExportMapping(
state, inst, uint_output ? EmitExportVec4U32(state, inst) : EmitExportVec4F32(state, inst),
vector_type);
if (inst.export_info.kind == IR::ExportTargetKind::Position) {
const auto pointer = state.builder.AllocateId();
state.builder.AddFunction({OpAccessChain, state.ptr_output_vec4_float, pointer, variable,
ConstantU32(state, 0)});
state.builder.AddFunction(
{OpAccessChain, state.ptr_output_vec4_float, pointer, variable, ConstantU32(state, 0)});
state.builder.AddFunction({OpStore, pointer, value});
return;
}
@@ -102,7 +102,7 @@ uint32_t EmitWqmLaneU32(EmitterState& state, uint32_t src) {
state.builder.AddFunction(
{OpINotEqual, state.bool_type, non_zero, masked, ConstantU32(state, 0)});
state.builder.AddFunction({OpSelect, state.uint_type, expanded, non_zero,
ConstantU32(state, mask), ConstantU32(state, 0)});
ConstantU32(state, mask), ConstantU32(state, 0)});
state.builder.AddFunction({OpBitwiseOr, state.uint_type, combined, ret, expanded});
ret = combined;
}
@@ -122,8 +122,8 @@ void EmitWqmB64(EmitterState& state, const IR::Instruction& inst) {
}
const auto ballot = state.builder.AllocateId();
state.builder.AddFunction({OpGroupNonUniformBallot, state.vec4_uint_type, ballot,
ConstantU32(state, ScopeSubgroup),
EmitLaneMaskOperandActiveBool(state, inst.src[0])});
ConstantU32(state, ScopeSubgroup),
EmitLaneMaskOperandActiveBool(state, inst.src[0])});
const auto low = state.builder.AllocateId();
const auto high = state.builder.AllocateId();
state.builder.AddFunction({OpCompositeExtract, state.uint_type, low, ballot, 0});
@@ -150,8 +150,8 @@ void EmitWqmB64(EmitterState& state, const IR::Instruction& inst) {
EmitPerInvocationMask(state, inst.dst, active);
} else {
const auto result = state.builder.AllocateId();
state.builder.AddFunction({OpSelect, state.uint_type, result, active,
ConstantU32(state, 1), ConstantU32(state, 0)});
state.builder.AddFunction({OpSelect, state.uint_type, result, active, ConstantU32(state, 1),
ConstantU32(state, 0)});
EmitStoreU32(state, inst.dst, result);
EmitStoreU32(state, OffsetRegisterOperand(inst.dst, 1), ConstantU32(state, 0));
}
@@ -205,8 +205,7 @@ void EmitSaveexecB32(EmitterState& state, const IR::Instruction& inst) {
const auto cond = state.builder.AllocateId();
const auto scc = state.builder.AllocateId();
state.builder.AddFunction(
{OpINotEqual, state.bool_type, cond, new_low, ConstantU32(state, 0)});
state.builder.AddFunction({OpINotEqual, state.bool_type, cond, new_low, ConstantU32(state, 0)});
state.builder.AddFunction(
{OpSelect, state.uint_type, scc, cond, ConstantU32(state, 1), ConstantU32(state, 0)});
EmitStoreU32(state, SccOperand(), scc);
@@ -269,18 +268,19 @@ void EmitReadFirstLaneU32(EmitterState& state, const IR::Instruction& inst) {
const auto first_lane = state.builder.AllocateId();
const auto first_value = state.builder.AllocateId();
state.builder.AddFunction({OpGroupNonUniformBallot, state.vec4_uint_type, ballot,
ConstantU32(state, ScopeSubgroup), active});
ConstantU32(state, ScopeSubgroup), active});
state.builder.AddFunction({OpGroupNonUniformBallotFindLSB, state.uint_type, first_lane,
ConstantU32(state, ScopeSubgroup), ballot});
ConstantU32(state, ScopeSubgroup), ballot});
state.builder.AddFunction({OpGroupNonUniformShuffle, state.uint_type, first_value,
ConstantU32(state, ScopeSubgroup), src, first_lane});
ConstantU32(state, ScopeSubgroup), src, first_lane});
EmitStoreU32(state, inst.dst, first_value);
}
uint32_t EmitLaneIndex(EmitterState& state, const IR::Operand& operand) {
const auto lane = state.builder.AllocateId();
const auto mask = state.wave_size == 32u ? 31u : 63u;
state.builder.AddFunction({OpBitwiseAnd, state.uint_type, lane, EmitValueLoad(state, operand),
ConstantU32(state, 63)});
ConstantU32(state, mask)});
return lane;
}
@@ -289,7 +289,7 @@ void EmitReadLaneU32(EmitterState& state, const IR::Instruction& inst) {
const auto lane = EmitLaneIndex(state, inst.src[1]);
const auto value = state.builder.AllocateId();
state.builder.AddFunction({OpGroupNonUniformShuffle, state.uint_type, value,
ConstantU32(state, ScopeSubgroup), src, lane});
ConstantU32(state, ScopeSubgroup), src, lane});
EmitStoreU32(state, inst.dst, value);
}
@@ -336,10 +336,8 @@ void EmitPermlaneB32(EmitterState& state, const IR::Instruction& inst, bool x16)
state.builder.AddFunction(
{OpBitwiseXor, state.uint_type, row_value, row, ConstantU32(state, 16)});
}
state.builder.AddFunction(
{OpBitwiseAnd, state.uint_type, lane, subid, ConstantU32(state, 15)});
state.builder.AddFunction(
{OpBitwiseAnd, state.uint_type, lane8, lane, ConstantU32(state, 7)});
state.builder.AddFunction({OpBitwiseAnd, state.uint_type, lane, subid, ConstantU32(state, 15)});
state.builder.AddFunction({OpBitwiseAnd, state.uint_type, lane8, lane, ConstantU32(state, 7)});
state.builder.AddFunction(
{OpShiftLeftLogical, state.uint_type, shift, lane8, ConstantU32(state, 2)});
state.builder.AddFunction(
@@ -350,7 +348,7 @@ void EmitPermlaneB32(EmitterState& state, const IR::Instruction& inst, bool x16)
{OpBitwiseAnd, state.uint_type, index1, index0, ConstantU32(state, 15)});
state.builder.AddFunction({OpBitwiseOr, state.uint_type, target, row_value, index1});
state.builder.AddFunction({OpGroupNonUniformShuffle, state.uint_type, shuffled,
ConstantU32(state, ScopeSubgroup), value, target});
ConstantU32(state, ScopeSubgroup), value, target});
uint32_t ret = shuffled;
if (!inst.dst.op_sel) {
const auto source_active = EmitLaneIndexActiveBool(state, target);
@@ -375,7 +373,7 @@ void EmitBarrier(EmitterState& state, const IR::Instruction& inst) {
(void)inst;
const auto semantics = MemorySemanticsAcquireRelease | MemorySemanticsWorkgroupMemory;
state.builder.AddFunction({OpControlBarrier, ConstantU32(state, ScopeWorkgroup),
ConstantU32(state, ScopeWorkgroup), ConstantU32(state, semantics)});
ConstantU32(state, ScopeWorkgroup), ConstantU32(state, semantics)});
}
} // namespace Libs::Graphics::ShaderRecompiler::Spirv::Emitter
@@ -2,6 +2,66 @@
#include "graphics/shader/recompiler/emitter/spirvEmitterInternal.h"
namespace Libs::Graphics::ShaderRecompiler::Spirv::Emitter {
namespace {
uint32_t EmitCubeAxisF32(EmitterState& state, uint32_t value) {
const auto normalized = state.builder.AllocateId();
state.builder.AddFunction(
{OpFSub, state.float_type, normalized, value, ConstantF32(state, 0x3f800000u)});
return normalized;
}
uint32_t EmitCubeLayerF32(EmitterState& state, uint32_t face_id) {
// Sampled RDNA2 cubemaps encode face_id as slice * 8 + face. The native
// 2D-array view stores six contiguous faces per slice, so remove the two
// reserved face IDs from every preceding slice.
const auto guest_layer = state.builder.AllocateId();
const auto slice = state.builder.AllocateId();
const auto padding = state.builder.AllocateId();
const auto host_layer = state.builder.AllocateId();
const auto result = state.builder.AllocateId();
state.builder.AddFunction({OpConvertFToU, state.uint_type, guest_layer, face_id});
state.builder.AddFunction(
{OpShiftRightLogical, state.uint_type, slice, guest_layer, ConstantU32(state, 3)});
state.builder.AddFunction(
{OpShiftLeftLogical, state.uint_type, padding, slice, ConstantU32(state, 1)});
state.builder.AddFunction({OpISub, state.uint_type, host_layer, guest_layer, padding});
state.builder.AddFunction({OpConvertUToF, state.float_type, result, host_layer});
return result;
}
uint32_t EmitImageCoordF32Impl(EmitterState& state, const IR::Instruction& inst,
const IR::Operand& address, uint32_t first_component,
uint32_t components) {
auto x = EmitImageAddressFloatLoad(state, inst, address, first_component);
if (components == 1u) {
return x;
}
auto y = inst.memory.image_address_components > first_component + 1u
? EmitImageAddressFloatLoad(state, inst, address, first_component + 1u)
: EmitZeroF32(state);
if (inst.memory.image_cube) {
// RDNA2 sampled cubemap S/T coordinates are biased by +1 relative to
// normalized 2D-array coordinates.
x = EmitCubeAxisF32(state, x);
y = EmitCubeAxisF32(state, y);
}
const auto coord = state.builder.AllocateId();
if (components == 3u) {
auto z = inst.memory.image_address_components > first_component + 2u
? EmitImageAddressFloatLoad(state, inst, address, first_component + 2u)
: EmitZeroF32(state);
if (inst.memory.image_cube) {
z = EmitCubeLayerF32(state, z);
}
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, coord, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, coord, x, y});
}
return coord;
}
} // namespace
bool HasImageSampleFlag(const IR::Instruction& inst, uint32_t flag) {
return (inst.memory.image_sample_flags & flag) != 0;
@@ -21,7 +81,7 @@ ImageSampleLayout MakeImageSampleLayout(const IR::Instruction& inst, ImageViewKi
}
if (HasImageSampleFlag(inst, Decoder::ImageSampleFlagDerivative)) {
const auto components = ImageViewSpatialComponents(view);
layout.grad_x = cursor;
layout.grad_x = cursor;
cursor += components;
layout.grad_y = cursor;
cursor += components;
@@ -36,24 +96,8 @@ ImageSampleLayout MakeImageSampleLayout(const IR::Instruction& inst, ImageViewKi
uint32_t EmitImageCoordF32(EmitterState& state, const IR::Instruction& inst,
const ImageSampleLayout& layout, ImageViewKind view) {
const auto x = EmitImageAddressFloatLoad(state, inst, inst.src[0], layout.coord);
const auto components = ImageViewCoordinateComponents(view);
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > layout.coord + 1u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], layout.coord + 1u)
: EmitZeroF32(state);
const auto coord = state.builder.AllocateId();
if (components == 3u) {
const auto z = inst.memory.image_address_components > layout.coord + 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], layout.coord + 2u)
: EmitZeroF32(state);
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, coord, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, coord, x, y});
}
return coord;
return EmitImageCoordF32Impl(state, inst, inst.src[0], layout.coord,
ImageViewCoordinateComponents(view));
}
uint32_t EmitImageLodF32(EmitterState& state, const IR::Instruction& inst,
@@ -95,10 +139,10 @@ uint32_t EmitImageGradientF32(EmitterState& state, const IR::Instruction& inst,
: EmitZeroF32(state);
const auto grad = state.builder.AllocateId();
if (components == 3u) {
const auto z = inst.memory.image_address_components > first_component + 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0],
first_component + 2u)
: EmitZeroF32(state);
const auto z =
inst.memory.image_address_components > first_component + 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], first_component + 2u)
: EmitZeroF32(state);
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, grad, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, grad, x, y});
@@ -120,8 +164,7 @@ uint32_t EmitImagePackedOffsetI32(EmitterState& state, const IR::Instruction& in
state.builder.AddFunction(
{OpCompositeConstruct, state.vec3_int_type, ret, zero, zero, zero});
} else {
state.builder.AddFunction(
{OpCompositeConstruct, state.vec2_int_type, ret, zero, zero});
state.builder.AddFunction({OpCompositeConstruct, state.vec2_int_type, ret, zero, zero});
}
return ret;
}
@@ -131,18 +174,18 @@ uint32_t EmitImagePackedOffsetI32(EmitterState& state, const IR::Instruction& in
const auto offset_x = state.builder.AllocateId();
state.builder.AddFunction({OpBitcast, state.int_type, packed_i32, packed_bits});
state.builder.AddFunction({OpBitFieldSExtract, state.int_type, offset_x, packed_i32,
ConstantI32(state, 0), ConstantI32(state, 6)});
ConstantI32(state, 0), ConstantI32(state, 6)});
if (components == 1u) {
return offset_x;
}
const auto offset_y = state.builder.AllocateId();
const auto offset = state.builder.AllocateId();
state.builder.AddFunction({OpBitFieldSExtract, state.int_type, offset_y, packed_i32,
ConstantI32(state, 8), ConstantI32(state, 6)});
ConstantI32(state, 8), ConstantI32(state, 6)});
if (components == 3u) {
const auto offset_z = state.builder.AllocateId();
state.builder.AddFunction({OpBitFieldSExtract, state.int_type, offset_z, packed_i32,
ConstantI32(state, 16), ConstantI32(state, 6)});
ConstantI32(state, 16), ConstantI32(state, 6)});
state.builder.AddFunction(
{OpCompositeConstruct, state.vec3_int_type, offset, offset_x, offset_y, offset_z});
} else {
@@ -158,7 +201,7 @@ uint32_t EmitImageCoordU32(EmitterState& state, const IR::Instruction& inst, Ima
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > 1u
const auto y = inst.memory.image_address_components > 1u
? EmitImageAddressValueLoad(state, inst, inst.src[1], 1)
: ConstantU32(state, 0);
const auto coord = state.builder.AllocateId();
@@ -180,7 +223,7 @@ uint32_t EmitImageLoadCoordU32(EmitterState& state, const IR::Instruction& inst,
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > 1u
const auto y = inst.memory.image_address_components > 1u
? EmitImageAddressValueLoad(state, inst, inst.src[0], 1)
: ConstantU32(state, 0);
const auto coord = state.builder.AllocateId();
@@ -209,24 +252,8 @@ uint32_t EmitImageMipLodU32(EmitterState& state, const IR::Instruction& inst,
uint32_t EmitImageQueryCoordF32(EmitterState& state, const IR::Instruction& inst,
ImageViewKind view) {
const auto x = EmitImageAddressFloatLoad(state, inst, inst.src[0], 0);
const auto components = ImageViewCoordinateComponents(view);
if (components == 1u) {
return x;
}
const auto y = inst.memory.image_address_components > 1u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], 1)
: EmitZeroF32(state);
const auto coord = state.builder.AllocateId();
if (components == 3u) {
const auto z = inst.memory.image_address_components > 2u
? EmitImageAddressFloatLoad(state, inst, inst.src[0], 2)
: EmitZeroF32(state);
state.builder.AddFunction({OpCompositeConstruct, state.vec3_float_type, coord, x, y, z});
} else {
state.builder.AddFunction({OpCompositeConstruct, state.vec2_float_type, coord, x, y});
}
return coord;
// OpImageQueryLod takes only the spatial coordinates, even for arrayed images.
return EmitImageCoordF32Impl(state, inst, inst.src[0], 0, ImageViewSpatialComponents(view));
}
uint32_t DmaskComponentIndex(uint32_t dmask, uint32_t component) {
@@ -35,9 +35,9 @@ uint32_t ConstantImageGatherHorizontalOffsets(EmitterState& state, ImageViewKind
uint32_t LoadStorageImageDescriptorAtIndex(EmitterState& state, uint32_t resource,
uint32_t array_index, bool uint_image,
ImageViewKind view) {
const auto kind = StorageBindingKind(uint_image, view);
const auto kind = StorageBindingKind(uint_image, view);
const auto& descriptors = state.storage_images[StorageImageIndex(uint_image, view)];
const auto pointer =
const auto pointer =
DescriptorElementPointer(state, descriptors.pointer_type, descriptors.variable, array_index,
kind, resource, "storage image descriptor array was not emitted");
const auto image = state.builder.AllocateId();
@@ -133,10 +133,18 @@ void EmitImageLoad(EmitterState& state, const IR::Instruction& inst) {
const bool integer = inst.memory.kind == IR::ResourceKind::ImageUint;
const auto color = state.builder.AllocateId();
state.builder.AddFunction({OpImageFetch, integer ? state.vec4_uint_type : state.vec4_float_type,
color, image, EmitImageLoadCoordU32(state, inst, view),
ImageOperandsLodMask,
EmitImageMipLodU32(state, inst, inst.src[0], view)});
const auto coord = EmitImageLoadCoordU32(state, inst, view);
if (ImageSpirvMultisampled(view) != 0) {
const auto sample = EmitImageAddressValueLoad(state, inst, inst.src[0],
ImageViewCoordinateComponents(view));
state.builder.AddFunction({OpImageFetch,
integer ? state.vec4_uint_type : state.vec4_float_type, color,
image, coord, ImageOperandsSampleMask, sample});
} else {
state.builder.AddFunction(
{OpImageFetch, integer ? state.vec4_uint_type : state.vec4_float_type, color, image,
coord, ImageOperandsLodMask, EmitImageMipLodU32(state, inst, inst.src[0], view)});
}
const auto dmask = inst.memory.dmask != 0 ? inst.memory.dmask : 1u;
uint32_t dst_index = 0;
@@ -158,8 +166,8 @@ void EmitImageLoad(EmitterState& state, const IR::Instruction& inst) {
void EmitImageStore(EmitterState& state, const IR::Instruction& inst) {
const auto uint_image = inst.memory.kind == IR::ResourceKind::StorageImageUint;
const auto view = StorageImageViewKind(state, inst.memory, uint_image, inst.pc);
const auto binding = ResourceForDescriptor(state, StorageBindingKind(uint_image, view),
inst.memory.resource);
const auto binding =
ResourceForDescriptor(state, StorageBindingKind(uint_image, view), inst.memory.resource);
const auto image = LoadStorageImageDescriptorAtIndex(state, inst.memory.resource,
binding.array_index, uint_image, view);
@@ -261,9 +269,9 @@ void EmitImageSample(EmitterState& state, const IR::Instruction& inst) {
} else if (integer) {
result_type = state.vec4_uint_type;
}
const auto explicit_lod = ImageSampleNeedsExplicitLod(state, inst);
const auto opcode = ImageSampleOpcode(state, inst);
std::vector<uint32_t> words = {opcode, result_type, sample, sampled_image, base_coord};
const auto explicit_lod = ImageSampleNeedsExplicitLod(state, inst);
const auto opcode = ImageSampleOpcode(state, inst);
std::vector<uint32_t> words = {opcode, result_type, sample, sampled_image, base_coord};
if (dref) {
words.push_back(EmitImageDrefF32(state, inst, layout));
}
@@ -99,6 +99,7 @@ enum : uint32_t {
ImageOperandsGradMask = 0x00000004u,
ImageOperandsOffsetMask = 0x00000010u,
ImageOperandsConstOffsetsMask = 0x00000020u,
ImageOperandsSampleMask = 0x00000040u,
};
enum : uint32_t {
@@ -150,7 +151,6 @@ enum : uint32_t {
OpImageGather = 96,
OpImageDrefGather = 97,
OpImageWrite = 99,
OpImage = 100,
OpImageQuerySizeLod = 103,
OpImageQueryLod = 105,
OpImageQueryLevels = 106,
@@ -382,7 +382,7 @@ struct EmitterState {
uint32_t ptr_workgroup_array = 0;
uint32_t ptr_workgroup_uint = 0;
uint32_t lds_variable = 0;
std::array<SampledImageDescriptors, 10> sampled_images;
std::array<SampledImageDescriptors, 14> sampled_images;
std::array<StorageImageDescriptors, 10> storage_images;
uint32_t sampler_type = 0;
uint32_t sampler_array_type = 0;
@@ -453,17 +453,20 @@ enum class ImageViewKind {
Dim2D,
Dim2DArray,
Dim3D,
Dim2DMsaa,
Dim2DMsaaArray,
Count,
};
constexpr uint32_t ImageViewKindCount = static_cast<uint32_t>(ImageViewKind::Count);
constexpr uint32_t SampledImageViewKindCount = static_cast<uint32_t>(ImageViewKind::Count);
constexpr uint32_t StorageImageViewKindCount = static_cast<uint32_t>(ImageViewKind::Dim2DMsaa);
constexpr uint32_t SampledImageIndex(bool integer, ImageViewKind view) {
return static_cast<uint32_t>(view) + (integer ? ImageViewKindCount : 0u);
return static_cast<uint32_t>(view) + (integer ? SampledImageViewKindCount : 0u);
}
constexpr uint32_t StorageImageIndex(bool integer, ImageViewKind view) {
return static_cast<uint32_t>(view) + (integer ? ImageViewKindCount : 0u);
return static_cast<uint32_t>(view) + (integer ? StorageImageViewKindCount : 0u);
}
constexpr IR::DescriptorBindingKind SampledBindingKind(bool integer, ImageViewKind view) {
@@ -474,6 +477,9 @@ constexpr IR::DescriptorBindingKind SampledBindingKind(bool integer, ImageViewKi
case ImageViewKind::Dim2D: return IR::DescriptorBindingKind::SampledUint2D;
case ImageViewKind::Dim2DArray: return IR::DescriptorBindingKind::SampledUint2DArray;
case ImageViewKind::Dim3D: return IR::DescriptorBindingKind::SampledUint3D;
case ImageViewKind::Dim2DMsaa: return IR::DescriptorBindingKind::SampledUint2DMsaa;
case ImageViewKind::Dim2DMsaaArray:
return IR::DescriptorBindingKind::SampledUint2DMsaaArray;
default: break;
}
}
@@ -483,6 +489,8 @@ constexpr IR::DescriptorBindingKind SampledBindingKind(bool integer, ImageViewKi
case ImageViewKind::Dim2D: return IR::DescriptorBindingKind::Sampled2D;
case ImageViewKind::Dim2DArray: return IR::DescriptorBindingKind::Sampled2DArray;
case ImageViewKind::Dim3D: return IR::DescriptorBindingKind::Sampled3D;
case ImageViewKind::Dim2DMsaa: return IR::DescriptorBindingKind::Sampled2DMsaa;
case ImageViewKind::Dim2DMsaaArray: return IR::DescriptorBindingKind::Sampled2DMsaaArray;
default: break;
}
return IR::DescriptorBindingKind::Count;
@@ -516,6 +524,8 @@ constexpr uint32_t ImageSpirvDimension(ImageViewKind view) {
case ImageViewKind::Dim1DArray: return Dim1D;
case ImageViewKind::Dim2D:
case ImageViewKind::Dim2DArray:
case ImageViewKind::Dim2DMsaa:
case ImageViewKind::Dim2DMsaaArray:
case ImageViewKind::Count: return Dim2D;
case ImageViewKind::Dim3D: return Dim3D;
}
@@ -523,7 +533,14 @@ constexpr uint32_t ImageSpirvDimension(ImageViewKind view) {
}
constexpr uint32_t ImageSpirvArrayed(ImageViewKind view) {
return view == ImageViewKind::Dim1DArray || view == ImageViewKind::Dim2DArray ? 1u : 0u;
return view == ImageViewKind::Dim1DArray || view == ImageViewKind::Dim2DArray ||
view == ImageViewKind::Dim2DMsaaArray
? 1u
: 0u;
}
constexpr uint32_t ImageSpirvMultisampled(ImageViewKind view) {
return view == ImageViewKind::Dim2DMsaa || view == ImageViewKind::Dim2DMsaaArray ? 1u : 0u;
}
struct AddCarryResult {
@@ -174,6 +174,12 @@ uint32_t VertexParameterInputPointerType(const EmitterState& state, VertexInputS
}
}
static bool MrtUsesUintOutput(const EmitterState& state, uint32_t index) {
return state.stage == ShaderType::Pixel && state.pixel_input_info != nullptr &&
index < std::size(state.pixel_input_info->target_output_mode) &&
state.pixel_input_info->target_output_mode[index] == 7u;
}
void AllocateInputVariables(EmitterState& state) {
for (auto& binding: state.inputs) {
binding.variable_id = state.builder.AllocateId();
@@ -323,23 +329,39 @@ void AddDescriptorAnnotationsAndNames(EmitterState& state) {
Decorate(state.address_memory_variable, "address_memory",
IR::DescriptorBindingKind::AddressMemory);
}
constexpr const char* SampledNames[] = {
"sampled_1d", "sampled_1d_array", "sampled_2d", "sampled_2d_array",
"sampled_3d", "sampled_uint_1d", "sampled_uint_1d_array",
"sampled_uint_2d", "sampled_uint_2d_array", "sampled_uint_3d"};
constexpr const char* SampledNames[] = {"sampled_1d",
"sampled_1d_array",
"sampled_2d",
"sampled_2d_array",
"sampled_3d",
"sampled_2d_msaa",
"sampled_2d_msaa_array",
"sampled_uint_1d",
"sampled_uint_1d_array",
"sampled_uint_2d",
"sampled_uint_2d_array",
"sampled_uint_3d",
"sampled_uint_2d_msaa",
"sampled_uint_2d_msaa_array"};
for (uint32_t i = 0; i < state.sampled_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const auto view = static_cast<ImageViewKind>(i % SampledImageViewKindCount);
Decorate(state.sampled_images[i].variable, SampledNames[i],
SampledBindingKind(i >= ImageViewKindCount, view));
SampledBindingKind(i >= SampledImageViewKindCount, view));
}
constexpr const char* StorageNames[] = {
"storage_1d", "storage_1d_array", "storage_2d", "storage_2d_array",
"storage_3d", "storage_uint_1d", "storage_uint_1d_array",
"storage_uint_2d", "storage_uint_2d_array", "storage_uint_3d"};
constexpr const char* StorageNames[] = {"storage_1d",
"storage_1d_array",
"storage_2d",
"storage_2d_array",
"storage_3d",
"storage_uint_1d",
"storage_uint_1d_array",
"storage_uint_2d",
"storage_uint_2d_array",
"storage_uint_3d"};
for (uint32_t i = 0; i < state.storage_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const auto view = static_cast<ImageViewKind>(i % StorageImageViewKindCount);
Decorate(state.storage_images[i].variable, StorageNames[i],
StorageBindingKind(i >= ImageViewKindCount, view));
StorageBindingKind(i >= StorageImageViewKindCount, view));
}
if (state.sampler_variable != 0) {
Decorate(state.sampler_variable, "samplers", IR::DescriptorBindingKind::Samplers);
@@ -409,6 +431,7 @@ void EmitHeaderAndTypes(EmitterState& state) {
state.ptr_output_sample_mask_array = state.builder.AllocateId();
state.ptr_output_float = state.builder.AllocateId();
state.ptr_output_vec4_float = state.builder.AllocateId();
const auto ptr_output_vec4_uint = state.builder.AllocateId();
state.per_vertex_type = state.builder.AllocateId();
state.ptr_output_per_vertex = state.builder.AllocateId();
state.storage_runtime_array_type = state.builder.AllocateId();
@@ -444,15 +467,15 @@ void EmitHeaderAndTypes(EmitterState& state) {
image.array_type = state.builder.AllocateId();
image.array_pointer_type = state.builder.AllocateId();
}
state.sampler_type = state.builder.AllocateId();
state.sampler_array_type = state.builder.AllocateId();
state.ptr_uniform_sampler = state.builder.AllocateId();
state.ptr_uniform_sampler_array = state.builder.AllocateId();
state.ptr_image_uint = state.builder.AllocateId();
state.func_type = state.builder.AllocateId();
state.main_func = state.builder.AllocateId();
state.entry_label = state.builder.AllocateId();
state.glsl_std450 = state.builder.AllocateId();
state.sampler_type = state.builder.AllocateId();
state.sampler_array_type = state.builder.AllocateId();
state.ptr_uniform_sampler = state.builder.AllocateId();
state.ptr_uniform_sampler_array = state.builder.AllocateId();
state.ptr_image_uint = state.builder.AllocateId();
state.func_type = state.builder.AllocateId();
state.main_func = state.builder.AllocateId();
state.entry_label = state.builder.AllocateId();
state.glsl_std450 = state.builder.AllocateId();
state.builder.AddCapability({CapabilityShader});
state.builder.AddCapability({CapabilitySampled1D});
@@ -462,7 +485,7 @@ void EmitHeaderAndTypes(EmitterState& state) {
state.builder.AddCapability({CapabilityImageGatherExtended});
}
if (std::any_of(state.storage_images.begin(),
state.storage_images.begin() + ImageViewKindCount,
state.storage_images.begin() + StorageImageViewKindCount,
[](const auto& image) { return image.variable != 0; })) {
state.builder.AddCapability({CapabilityStorageImageReadWithoutFormat});
state.builder.AddCapability({CapabilityStorageImageWriteWithoutFormat});
@@ -605,6 +628,8 @@ void EmitHeaderAndTypes(EmitterState& state) {
{OpTypePointer, state.ptr_output_int, StorageClassOutput, state.int_type});
state.builder.AddType(
{OpTypePointer, state.ptr_output_vec4_float, StorageClassOutput, state.vec4_float_type});
state.builder.AddType(
{OpTypePointer, ptr_output_vec4_uint, StorageClassOutput, state.vec4_uint_type});
if (state.per_vertex_variable != 0) {
state.builder.AddType({OpTypeStruct, state.per_vertex_type, state.vec4_float_type});
state.builder.AddType({OpTypePointer, state.ptr_output_per_vertex, StorageClassOutput,
@@ -615,8 +640,12 @@ void EmitHeaderAndTypes(EmitterState& state) {
for (const auto& binding: state.outputs) {
if (binding.kind == IR::StageOutputKind::Parameter ||
binding.kind == IR::StageOutputKind::Mrt) {
const auto pointer_type =
binding.kind == IR::StageOutputKind::Mrt && MrtUsesUintOutput(state, binding.index)
? ptr_output_vec4_uint
: state.ptr_output_vec4_float;
state.builder.AddType(
{OpVariable, state.ptr_output_vec4_float, binding.variable_id, StorageClassOutput});
{OpVariable, pointer_type, binding.variable_id, StorageClassOutput});
}
}
if (state.depth_variable != 0) {
@@ -700,11 +729,11 @@ void EmitHeaderAndTypes(EmitterState& state) {
}
for (uint32_t i = 0; i < state.sampled_images.size(); i++) {
auto& image = state.sampled_images[i];
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const bool integer = i >= ImageViewKindCount;
const auto view = static_cast<ImageViewKind>(i % SampledImageViewKindCount);
const bool integer = i >= SampledImageViewKindCount;
const auto component = integer ? state.uint_type : state.float_type;
state.builder.AddType({OpTypeImage, image.image_type, component,
ImageSpirvDimension(view), 0, ImageSpirvArrayed(view), 0, 1,
state.builder.AddType({OpTypeImage, image.image_type, component, ImageSpirvDimension(view),
0, ImageSpirvArrayed(view), ImageSpirvMultisampled(view), 1,
ImageFormatUnknown});
state.builder.AddType({OpTypeSampledImage, image.sampled_image_type, image.image_type});
state.builder.AddType(
@@ -733,13 +762,12 @@ void EmitHeaderAndTypes(EmitterState& state) {
}
for (uint32_t i = 0; i < state.storage_images.size(); i++) {
auto& image = state.storage_images[i];
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
const bool integer = i >= ImageViewKindCount;
const auto view = static_cast<ImageViewKind>(i % StorageImageViewKindCount);
const bool integer = i >= StorageImageViewKindCount;
const auto component = integer ? state.uint_type : state.float_type;
const auto format = integer ? ImageFormatR32ui : ImageFormatUnknown;
state.builder.AddType({OpTypeImage, image.image_type, component,
ImageSpirvDimension(view), 0, ImageSpirvArrayed(view), 0, 2,
format});
state.builder.AddType({OpTypeImage, image.image_type, component, ImageSpirvDimension(view),
0, ImageSpirvArrayed(view), 0, 2, format});
state.builder.AddType(
{OpTypePointer, image.pointer_type, StorageClassUniformConstant, image.image_type});
if (image.variable != 0) {
@@ -786,15 +814,15 @@ void AllocateDescriptorVariables(EmitterState& state) {
state.flattened_srt_variable = state.builder.AllocateId();
}
for (uint32_t i = 0; i < state.sampled_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
if (DescriptorBinding(state, SampledBindingKind(i >= ImageViewKindCount, view)) !=
const auto view = static_cast<ImageViewKind>(i % SampledImageViewKindCount);
if (DescriptorBinding(state, SampledBindingKind(i >= SampledImageViewKindCount, view)) !=
nullptr) {
state.sampled_images[i].variable = state.builder.AllocateId();
}
}
for (uint32_t i = 0; i < state.storage_images.size(); i++) {
const auto view = static_cast<ImageViewKind>(i % ImageViewKindCount);
if (DescriptorBinding(state, StorageBindingKind(i >= ImageViewKindCount, view)) !=
const auto view = static_cast<ImageViewKind>(i % StorageImageViewKindCount);
if (DescriptorBinding(state, StorageBindingKind(i >= StorageImageViewKindCount, view)) !=
nullptr) {
state.storage_images[i].variable = state.builder.AllocateId();
}
@@ -13,16 +13,30 @@ namespace {
constexpr uint32_t MaxPushConstantBytes = 128;
constexpr std::array ImageBindingKinds = {
DescriptorBindingKind::Sampled1D, DescriptorBindingKind::Sampled1DArray,
DescriptorBindingKind::Sampled2D, DescriptorBindingKind::Sampled2DArray,
DescriptorBindingKind::Sampled3D, DescriptorBindingKind::SampledUint1D,
DescriptorBindingKind::SampledUint1DArray, DescriptorBindingKind::SampledUint2D,
DescriptorBindingKind::SampledUint2DArray, DescriptorBindingKind::SampledUint3D,
DescriptorBindingKind::Storage1D, DescriptorBindingKind::Storage1DArray,
DescriptorBindingKind::Storage2D, DescriptorBindingKind::Storage2DArray,
DescriptorBindingKind::Storage3D, DescriptorBindingKind::StorageUint1D,
DescriptorBindingKind::StorageUint1DArray, DescriptorBindingKind::StorageUint2D,
DescriptorBindingKind::StorageUint2DArray, DescriptorBindingKind::StorageUint3D,
DescriptorBindingKind::Sampled1D,
DescriptorBindingKind::Sampled1DArray,
DescriptorBindingKind::Sampled2D,
DescriptorBindingKind::Sampled2DArray,
DescriptorBindingKind::Sampled2DMsaa,
DescriptorBindingKind::Sampled2DMsaaArray,
DescriptorBindingKind::Sampled3D,
DescriptorBindingKind::SampledUint1D,
DescriptorBindingKind::SampledUint1DArray,
DescriptorBindingKind::SampledUint2D,
DescriptorBindingKind::SampledUint2DArray,
DescriptorBindingKind::SampledUint2DMsaa,
DescriptorBindingKind::SampledUint2DMsaaArray,
DescriptorBindingKind::SampledUint3D,
DescriptorBindingKind::Storage1D,
DescriptorBindingKind::Storage1DArray,
DescriptorBindingKind::Storage2D,
DescriptorBindingKind::Storage2DArray,
DescriptorBindingKind::Storage3D,
DescriptorBindingKind::StorageUint1D,
DescriptorBindingKind::StorageUint1DArray,
DescriptorBindingKind::StorageUint2D,
DescriptorBindingKind::StorageUint2DArray,
DescriptorBindingKind::StorageUint3D,
};
bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
@@ -36,6 +50,8 @@ bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
case Dimension::Dim1DArray: result = Kind::Sampled1DArray; return true;
case Dimension::Dim2D: result = Kind::Sampled2D; return true;
case Dimension::Dim2DArray: result = Kind::Sampled2DArray; return true;
case Dimension::Dim2DMsaa: result = Kind::Sampled2DMsaa; return true;
case Dimension::Dim2DMsaaArray: result = Kind::Sampled2DMsaaArray; return true;
case Dimension::Dim3D: result = Kind::Sampled3D; return true;
default: return false;
}
@@ -45,6 +61,8 @@ bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
case Dimension::Dim1DArray: result = Kind::SampledUint1DArray; return true;
case Dimension::Dim2D: result = Kind::SampledUint2D; return true;
case Dimension::Dim2DArray: result = Kind::SampledUint2DArray; return true;
case Dimension::Dim2DMsaa: result = Kind::SampledUint2DMsaa; return true;
case Dimension::Dim2DMsaaArray: result = Kind::SampledUint2DMsaaArray; return true;
case Dimension::Dim3D: result = Kind::SampledUint3D; return true;
default: return false;
}
@@ -71,7 +89,7 @@ bool ImageBinding(const ImageResource& image, DescriptorBindingKind& result) {
}
bool CollectValue(const ScalarProvenance& provenance, uint32_t id, std::vector<uint8_t>& visited,
std::set<uint32_t>& registers) {
std::set<uint32_t>& registers) {
if (id <= ScalarProvenance::Unknown) {
return true;
}
@@ -104,7 +122,7 @@ bool CollectValue(const ScalarProvenance& provenance, uint32_t id, std::vector<u
}
bool CollectSource(const Program& program, uint32_t source, bool allow_unknown,
std::vector<uint8_t>& visited, std::set<uint32_t>& registers) {
std::vector<uint8_t>& visited, std::set<uint32_t>& registers) {
if (allow_unknown && source == ScalarProvenance::Unknown) {
return true;
}
@@ -165,8 +183,7 @@ bool CollectUserData(const Program& program, std::vector<uint32_t>& result) {
return false;
}
for (uint32_t i = 0; i < inst.src_count; i++) {
if (!CollectValue(program.provenance, inst.scalar_sources[i], visited,
registers)) {
if (!CollectValue(program.provenance, inst.scalar_sources[i], visited, registers)) {
return false;
}
}
@@ -199,7 +216,7 @@ bool AllocateBindings(Program& program, const BindingLayoutOptions& options, std
if (!program.shader_info_complete || program.binding_layout_complete) {
if (error != nullptr) {
*error = !program.shader_info_complete ? "shader info is not ready"
: "binding layout already allocated";
: "binding layout already allocated";
}
return false;
}
@@ -0,0 +1,323 @@
#include "graphics/shader/recompiler/ir/ReadLaneElimination.h"
#include "graphics/shader/recompiler/ir/SrtWalker.h"
#include <algorithm>
#include <iterator>
#include <map>
#include <set>
#include <utility>
namespace Libs::Graphics::ShaderRecompiler::IR {
namespace {
constexpr uint32_t FirstTemporaryScalarRegister = 128;
struct LaneKey {
uint32_t reg = 0;
uint32_t lane = 0;
auto operator<=>(const LaneKey&) const = default;
};
using LaneSet = std::set<LaneKey>;
bool PairDwordOpcode(Opcode op) {
switch (op) {
case Opcode::MoveU64:
case Opcode::WqmB64:
case Opcode::SaveexecB64:
case Opcode::BitwiseAndU64:
case Opcode::BitwiseAndNotU64:
case Opcode::BitwiseOrU64:
case Opcode::BitwiseOrNotU64:
case Opcode::BitwiseXorU64:
case Opcode::BitwiseNandU64:
case Opcode::BitwiseNorU64:
case Opcode::BitwiseXnorU64:
case Opcode::BitwiseNotU64:
case Opcode::BitFieldMaskU64:
case Opcode::BitFieldExtractU64:
case Opcode::BitReplicateB64B32:
case Opcode::ShiftLeftLogicalU64:
case Opcode::ShiftRightLogicalU64:
case Opcode::SelectU64: return true;
default: return false;
}
}
bool ResolveLane(const Program& program, const Instruction& inst, uint32_t source_index,
uint32_t& lane) {
if (source_index >= inst.src_count || (program.wave_size != 32 && program.wave_size != 64)) {
return false;
}
const auto& selector = inst.src[source_index];
if (selector.kind == OperandKind::ImmediateU32) {
lane = selector.imm % program.wave_size;
return true;
}
uint32_t folded = 0;
if (!FoldScalarConstant(program.provenance, inst.scalar_sources[source_index], folded)) {
return false;
}
lane = folded % program.wave_size;
return true;
}
bool UniformWriteSource(const Instruction& inst) {
if (inst.src_count == 0) {
return false;
}
const auto& source = inst.src[0];
if (source.kind == OperandKind::ImmediateU32 || source.kind == OperandKind::PcRelativeU32) {
return true;
}
return source.kind == OperandKind::Register &&
(source.reg.file == RegisterFile::Scalar || source.reg.file == RegisterFile::Scc ||
source.reg.file == RegisterFile::M0);
}
bool WriteLaneKey(const Program& program, const Instruction& inst, LaneKey& key) {
if (inst.op != Opcode::WriteLaneU32 || inst.dst.kind != OperandKind::Register ||
inst.dst.reg.file != RegisterFile::Vector || !UniformWriteSource(inst)) {
return false;
}
uint32_t lane = 0;
if (!ResolveLane(program, inst, 1, lane)) {
return false;
}
key = {inst.dst.reg.index, lane};
return true;
}
bool ReadLaneKey(const Program& program, const Instruction& inst, LaneKey& key) {
if (inst.op != Opcode::ReadLaneU32 || inst.src_count < 2 ||
inst.src[0].kind != OperandKind::Register || inst.src[0].reg.file != RegisterFile::Vector) {
return false;
}
uint32_t lane = 0;
if (!ResolveLane(program, inst, 1, lane)) {
return false;
}
key = {inst.src[0].reg.index, lane};
return true;
}
void InvalidateRegister(LaneSet& valid, uint32_t reg) {
const auto first = valid.lower_bound({reg, 0});
const auto last = valid.lower_bound({reg + 1u, 0});
valid.erase(first, last);
}
void ApplyInstruction(const Program& program, const Instruction& inst, LaneSet& valid) {
if (inst.op == Opcode::WriteLaneU32 && inst.dst.kind == OperandKind::Register &&
inst.dst.reg.file == RegisterFile::Vector) {
LaneKey key;
if (WriteLaneKey(program, inst, key)) {
valid.insert(key);
return;
}
uint32_t lane = 0;
if (ResolveLane(program, inst, 1, lane)) {
valid.erase({inst.dst.reg.index, lane});
} else {
InvalidateRegister(valid, inst.dst.reg.index);
}
return;
}
if (inst.op == Opcode::MoveRelDestU32 && inst.dst.kind == OperandKind::Register &&
inst.dst.reg.file == RegisterFile::Vector) {
valid.clear();
return;
}
if (inst.dst.kind == OperandKind::Register && inst.dst.reg.file == RegisterFile::Vector) {
uint32_t dwords = std::max(inst.memory.data_dwords, 1u);
if (PairDwordOpcode(inst.op) || inst.op == Opcode::UMadU64U32) {
dwords = std::max(dwords, 2u);
}
for (uint32_t i = 0; i < dwords && inst.dst.reg.index <= UINT32_MAX - i; i++) {
InvalidateRegister(valid, inst.dst.reg.index + i);
}
}
if (inst.dst2.kind == OperandKind::Register && inst.dst2.reg.file == RegisterFile::Vector) {
InvalidateRegister(valid, inst.dst2.reg.index);
}
}
LaneSet TransferBlock(const Program& program, const BasicBlock& block, LaneSet state) {
for (const auto& inst: block.instructions) {
ApplyInstruction(program, inst, state);
}
return state;
}
LaneSet Intersect(const LaneSet& left, const LaneSet& right) {
LaneSet result;
std::set_intersection(left.begin(), left.end(), right.begin(), right.end(),
std::inserter(result, result.end()));
return result;
}
uint32_t NextTemporaryScalarRegister(const Program& program) {
uint32_t next = FirstTemporaryScalarRegister;
const auto consider = [&next](const Operand& operand) {
if (operand.kind == OperandKind::Register && operand.reg.file == RegisterFile::Scalar &&
operand.reg.index >= next && operand.reg.index != UINT32_MAX) {
next = operand.reg.index + 1u;
}
};
for (const auto& block: program.blocks) {
for (const auto& inst: block.instructions) {
consider(inst.dst);
consider(inst.dst2);
for (uint32_t i = 0; i < inst.src_count; i++) {
consider(inst.src[i]);
}
}
}
return next;
}
Operand ScalarRegisterOperand(uint32_t reg) {
Operand operand;
operand.kind = OperandKind::Register;
operand.reg.file = RegisterFile::Scalar;
operand.reg.index = reg;
return operand;
}
Instruction ShadowWrite(const Instruction& write, uint32_t temporary) {
Instruction shadow;
shadow.pc = write.pc;
shadow.op = Opcode::MoveU32;
shadow.dst = ScalarRegisterOperand(temporary);
shadow.src[0] = write.src[0];
shadow.src_count = 1;
return shadow;
}
Instruction ShadowRead(const Instruction& read, uint32_t temporary) {
Instruction rewritten;
rewritten.pc = read.pc;
rewritten.op = Opcode::MoveU32;
rewritten.dst = read.dst;
rewritten.src[0] = ScalarRegisterOperand(temporary);
rewritten.src_count = 1;
return rewritten;
}
} // namespace
ReadLaneEliminationStats EliminateReadLane(Program& program) {
ReadLaneEliminationStats stats;
if (program.blocks.empty() || (program.wave_size != 32 && program.wave_size != 64)) {
return stats;
}
LaneSet universe;
for (const auto& block: program.blocks) {
for (const auto& inst: block.instructions) {
LaneKey key;
if (WriteLaneKey(program, inst, key)) {
universe.insert(key);
}
}
}
if (universe.empty()) {
return stats;
}
const size_t block_count = program.blocks.size();
std::vector<LaneSet> entry(block_count, universe);
std::vector<LaneSet> exit(block_count, universe);
entry[0].clear();
for (size_t block = 0; block < block_count; block++) {
exit[block] = TransferBlock(program, program.blocks[block], entry[block]);
}
bool changed = true;
while (changed) {
changed = false;
for (size_t block_index = 0; block_index < block_count; block_index++) {
LaneSet next_entry;
const auto& block = program.blocks[block_index];
if (block_index != 0 && !block.predecessors.empty()) {
next_entry = universe;
for (const auto predecessor: block.predecessors) {
if (predecessor >= block_count) {
next_entry.clear();
break;
}
next_entry = Intersect(next_entry, exit[predecessor]);
}
}
auto next_exit = TransferBlock(program, block, next_entry);
if (next_entry != entry[block_index] || next_exit != exit[block_index]) {
entry[block_index] = std::move(next_entry);
exit[block_index] = std::move(next_exit);
changed = true;
}
}
}
LaneSet forwarded;
for (size_t block_index = 0; block_index < block_count; block_index++) {
auto state = entry[block_index];
for (const auto& inst: program.blocks[block_index].instructions) {
LaneKey key;
if (ReadLaneKey(program, inst, key) && state.contains(key)) {
forwarded.insert(key);
}
ApplyInstruction(program, inst, state);
}
}
if (forwarded.empty()) {
return stats;
}
std::map<LaneKey, uint32_t> temporaries;
auto next_temporary = NextTemporaryScalarRegister(program);
for (const auto& key: forwarded) {
if (next_temporary == UINT32_MAX) {
return {};
}
temporaries.emplace(key, next_temporary++);
}
for (size_t block_index = 0; block_index < block_count; block_index++) {
const auto original = std::move(program.blocks[block_index].instructions);
auto& rewritten = program.blocks[block_index].instructions;
rewritten.clear();
rewritten.reserve(original.size() + temporaries.size());
auto state = entry[block_index];
for (const auto& inst: original) {
LaneKey read_key;
if (ReadLaneKey(program, inst, read_key) && state.contains(read_key)) {
const auto temporary = temporaries.find(read_key);
if (temporary != temporaries.end()) {
rewritten.push_back(ShadowRead(inst, temporary->second));
stats.rewritten_reads++;
ApplyInstruction(program, inst, state);
continue;
}
}
rewritten.push_back(inst);
LaneKey write_key;
if (WriteLaneKey(program, inst, write_key)) {
const auto temporary = temporaries.find(write_key);
if (temporary != temporaries.end()) {
rewritten.push_back(ShadowWrite(inst, temporary->second));
stats.shadow_writes++;
}
}
ApplyInstruction(program, inst, state);
}
}
return stats;
}
} // namespace Libs::Graphics::ShaderRecompiler::IR
@@ -0,0 +1,20 @@
#ifndef EMULATOR_INCLUDE_EMULATOR_GRAPHICS_SHADER_RECOMPILER_READLANEELIMINATION_H_
#define EMULATOR_INCLUDE_EMULATOR_GRAPHICS_SHADER_RECOMPILER_READLANEELIMINATION_H_
#include "graphics/shader/recompiler/ir/ShaderIR.h"
namespace Libs::Graphics::ShaderRecompiler::IR {
struct ReadLaneEliminationStats {
uint32_t rewritten_reads = 0;
uint32_t shadow_writes = 0;
};
// Replaces fixed-lane ReadLane operations that are reached by a matching WriteLane on every
// control-flow path. A synthetic scalar register snapshots the value at WriteLane execution time,
// so the rewrite remains valid when the source SGPR is subsequently overwritten.
[[nodiscard]] ReadLaneEliminationStats EliminateReadLane(Program& program);
} // namespace Libs::Graphics::ShaderRecompiler::IR
#endif /* EMULATOR_INCLUDE_EMULATOR_GRAPHICS_SHADER_RECOMPILER_READLANEELIMINATION_H_ */
@@ -12,10 +12,11 @@ namespace {
constexpr uint64_t AddressMask = 0x0000ffffffffffffull;
Decoder::ImageDimension DescriptorDimension(const DescriptorValue& descriptor,
Decoder::ImageDimension DescriptorDimension(const DescriptorValue& descriptor,
Decoder::ImageDimension requested) {
const bool is_array = requested == Decoder::ImageDimension::Dim1DArray ||
requested == Decoder::ImageDimension::Dim2DArray;
requested == Decoder::ImageDimension::Dim2DArray ||
requested == Decoder::ImageDimension::Dim2DMsaaArray;
switch (static_cast<Prospero::ImageType>((descriptor.dwords[3] >> 28u) & 0xfu)) {
case Prospero::ImageType::kColor1D: return Decoder::ImageDimension::Dim1D;
case Prospero::ImageType::kColor1DArray:
@@ -26,13 +27,17 @@ Decoder::ImageDimension DescriptorDimension(const DescriptorValue& descrip
case Prospero::ImageType::kColor3D: return Decoder::ImageDimension::Dim3D;
case Prospero::ImageType::kCube: return Decoder::ImageDimension::Dim2DArray;
case Prospero::ImageType::kColor2DArray:
case Prospero::ImageType::kColor2DMsaaArray:
if (is_array) {
return Decoder::ImageDimension::Dim2DArray;
}
return Decoder::ImageDimension::Dim2D;
case Prospero::ImageType::kColor2D:
case Prospero::ImageType::kColor2DMsaa: return Decoder::ImageDimension::Dim2D;
case Prospero::ImageType::kColor2DMsaaArray:
if (is_array) {
return Decoder::ImageDimension::Dim2DMsaaArray;
}
return Decoder::ImageDimension::Dim2DMsaa;
case Prospero::ImageType::kColor2D: return Decoder::ImageDimension::Dim2D;
case Prospero::ImageType::kColor2DMsaa: return Decoder::ImageDimension::Dim2DMsaa;
default: return Decoder::ImageDimension::Unknown;
}
}
@@ -51,8 +56,7 @@ bool ValidImageDescriptor(const DescriptorValue& descriptor) {
const auto base_level = (descriptor.dwords[3] >> 12u) & 0xfu;
const auto fragments = (descriptor.dwords[3] >> 16u) & 0xfu;
const auto max_mip = (descriptor.dwords[5] >> 4u) & 0xfu;
return base_level == 0 && fragments >= 1 && fragments <= 3 &&
max_mip == fragments;
return base_level == 0 && fragments >= 1 && fragments <= 3 && max_mip == fragments;
}
return true;
}
@@ -61,6 +65,11 @@ uint32_t DescriptorImageSwizzle(const DescriptorValue& descriptor) {
return descriptor.dwords[3] & 0xfffu;
}
bool DescriptorIsCube(const DescriptorValue& descriptor) {
return static_cast<Prospero::ImageType>((descriptor.dwords[3] >> 28u) & 0xfu) ==
Prospero::ImageType::kCube;
}
bool DecodeBufferDescriptor(const DescriptorValue& descriptor, ShaderBufferResource& result) {
if (descriptor.dword_count != std::size(result.fields)) {
return false;
@@ -171,12 +180,13 @@ bool ValidateResourceSpecialization(const Program& program, const ResourceSnapsh
const auto& image = program.info.images[i];
const auto& descriptor = snapshot.images[i];
if (NullImageDescriptor(descriptor)) {
bool canonical_kind = image.kind == ResourceKind::Image ||
image.kind == ResourceKind::StorageImage;
bool canonical_kind =
image.kind == ResourceKind::Image || image.kind == ResourceKind::StorageImage;
if (image.atomic) {
canonical_kind = image.kind == ResourceKind::StorageImageUint;
}
if (image.dimension != Decoder::ImageDimension::Dim2D || !canonical_kind) {
if (image.dimension != Decoder::ImageDimension::Dim2D || image.cube ||
!canonical_kind) {
if (error != nullptr) {
*error = fmt::format(
"image descriptor {} no longer matches canonical null specialization", i);
@@ -186,7 +196,8 @@ bool ValidateResourceSpecialization(const Program& program, const ResourceSnapsh
continue;
}
const auto dimension = DescriptorDimension(descriptor, image.dimension);
if (dimension == Decoder::ImageDimension::Unknown || dimension != image.dimension) {
if (dimension == Decoder::ImageDimension::Unknown || dimension != image.dimension ||
DescriptorIsCube(descriptor) != image.cube) {
if (error != nullptr) {
*error =
fmt::format("image descriptor {} no longer matches specialized dimension", i);
@@ -361,6 +372,7 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
auto& image = next.images[i];
if (NullImageDescriptor(descriptor)) {
image.dimension = Decoder::ImageDimension::Dim2D;
image.cube = false;
switch (image.kind) {
case ResourceKind::ImageUint: image.kind = ResourceKind::Image; break;
case ResourceKind::StorageImageUint:
@@ -386,6 +398,7 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
return false;
}
image.dimension = descriptor_dimension;
image.cube = DescriptorIsCube(descriptor);
if (image.kind == ResourceKind::StorageImage ||
image.kind == ResourceKind::StorageImageUint) {
image.storage_swizzle = DescriptorImageSwizzle(descriptor);
@@ -402,6 +415,7 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
std::reference_wrapper<Instruction> inst;
ResourceKind kind;
Decoder::ImageDimension dimension;
bool cube;
};
std::vector<ImagePatch> patches;
for (auto& block: program.blocks) {
@@ -420,13 +434,14 @@ bool SpecializeResources(Program& program, const ResourceSnapshot& snapshot, std
return false;
}
const auto& image = next.images[inst.memory.resource];
patches.push_back({std::ref(inst), image.kind, image.dimension});
patches.push_back({std::ref(inst), image.kind, image.dimension, image.cube});
}
}
program.info = std::move(next);
for (const auto& patch: patches) {
patch.inst.get().memory.kind = patch.kind;
patch.inst.get().memory.image_dimension = patch.dimension;
patch.inst.get().memory.image_cube = patch.cube;
}
return true;
}
@@ -31,7 +31,7 @@ bool ValidateResourceSpecialization(const Program& program, const ResourceSnapsh
// Resolves the immutable dense resource topology against one runtime user-data/SRT snapshot.
// On failure the destination is unchanged.
bool MaterializeResources(const Program& program, const SrtRuntime& runtime,
ResourceSnapshot& snapshot, std::string* error);
ResourceSnapshot& snapshot, std::string* error);
// Applies runtime descriptor shape/format facts to a copied dense topology before layout and
// emission. On failure the program is unchanged.
@@ -84,7 +84,7 @@ uint32_t ByteExtent(const Instruction& inst) {
}
bool ContainsUnknown(const ScalarProvenance& provenance, uint32_t id, std::vector<uint8_t>& visited,
std::vector<uint32_t>& path) {
std::vector<uint32_t>& path) {
path.push_back(id);
if (id <= ScalarProvenance::Unknown || id >= provenance.values.size()) {
return true;
@@ -115,7 +115,7 @@ bool ContainsUnknown(const ScalarProvenance& provenance, uint32_t id, std::vecto
}
bool IsLoopInvariantValue(const ScalarProvenance& provenance, uint32_t id,
std::vector<uint8_t>& visiting) {
std::vector<uint8_t>& visiting) {
if (id <= ScalarProvenance::Unknown || id >= provenance.values.size()) {
return false;
}
@@ -432,6 +432,7 @@ struct MemoryInfo {
bool typed = false;
bool formatted = false;
bool image_has_mip = false;
bool image_cube = false;
bool glc = false;
bool slc = false;
bool idxen = false;
@@ -607,6 +608,7 @@ struct ImageResource {
bool written = false;
bool atomic = false;
bool depth_compare = false;
bool cube = false;
bool operator==(const ImageResource& other) const = default;
};
@@ -677,11 +679,15 @@ enum class DescriptorBindingKind {
Sampled1DArray,
Sampled2D,
Sampled2DArray,
Sampled2DMsaa,
Sampled2DMsaaArray,
Sampled3D,
SampledUint1D,
SampledUint1DArray,
SampledUint2D,
SampledUint2DArray,
SampledUint2DMsaa,
SampledUint2DMsaaArray,
SampledUint3D,
Storage1D,
Storage1DArray,
+14 -14
View File
@@ -530,8 +530,8 @@ bool BuildSrtPlan(Program& program, std::string* error) {
}
bool EvaluateDescriptorSource(const Program& program, uint32_t source, uint32_t use_pc,
const SrtRuntime& runtime, DescriptorValue& result,
std::string* error) {
const SrtRuntime& runtime, DescriptorValue& result,
std::string* error) {
const DescriptorSourceRequest request {source, use_pc};
std::vector<DescriptorValue> results;
if (!EvaluateDescriptorSources(program, std::span {&request, 1}, runtime, results, error)) {
@@ -542,11 +542,11 @@ bool EvaluateDescriptorSource(const Program& program, uint32_t source, uint32_t
}
static bool EvaluateRuntimeSourcesImpl(const Program& program,
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime,
std::vector<DescriptorValue>& results,
std::vector<uint32_t>& flat, bool evaluate_flat,
std::string* error) {
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime,
std::vector<DescriptorValue>& results,
std::vector<uint32_t>& flat, bool evaluate_flat,
std::string* error) {
if (!program.srt_plan_complete) {
if (error != nullptr) {
*error = Diagnostic(program, 0, "SRT plan is not ready");
@@ -602,22 +602,22 @@ static bool EvaluateRuntimeSourcesImpl(const Program&
}
bool EvaluateDescriptorSources(const Program& program,
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::string* error) {
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::string* error) {
std::vector<uint32_t> ignored;
return EvaluateRuntimeSourcesImpl(program, requests, runtime, results, ignored, false, error);
}
bool EvaluateRuntimeSources(const Program& program,
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::vector<uint32_t>& flat, std::string* error) {
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::vector<uint32_t>& flat, std::string* error) {
return EvaluateRuntimeSourcesImpl(program, requests, runtime, results, flat, true, error);
}
bool WalkSrt(const Program& program, const SrtRuntime& runtime, std::vector<uint32_t>& flat,
std::string* error) {
std::string* error) {
std::vector<DescriptorValue> ignored;
return EvaluateRuntimeSources(program, {}, runtime, ignored, flat, error);
}
@@ -30,22 +30,22 @@ bool FoldScalarConstant(const ScalarProvenance& provenance, uint32_t value, uint
bool BuildSrtPlan(Program& program, std::string* error);
bool EvaluateDescriptorSource(const Program& program, uint32_t source, uint32_t use_pc,
const SrtRuntime& runtime, DescriptorValue& result,
const SrtRuntime& runtime, DescriptorValue& result,
std::string* error);
// Evaluates one runtime snapshot transactionally. Scalar values and ReadConst results shared by
// several descriptors are memoized once across the batch.
bool EvaluateDescriptorSources(const Program& program,
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::string* error);
// Evaluates descriptor sources and the flattened immediate SRT with one memoized scalar walk.
// On failure neither destination is changed.
bool EvaluateRuntimeSources(const Program& program,
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::vector<uint32_t>& flat, std::string* error);
std::span<const DescriptorSourceRequest> requests,
const SrtRuntime& runtime, std::vector<DescriptorValue>& results,
std::vector<uint32_t>& flat, std::string* error);
bool WalkSrt(const Program& program, const SrtRuntime& runtime, std::vector<uint32_t>& flat,
std::string* error);
+9 -12
View File
@@ -13,8 +13,8 @@
#include "graphics/guest_gpu/graphicsRun.h"
#include "graphics/guest_gpu/hardwareContext.h"
#include "graphics/host_gpu/renderer/renderContext.h"
#include "graphics/shader/recompiler/decompiler/ShaderDecoder.h"
#include "graphics/shader/recompiler/ShaderRecompiler.h"
#include "graphics/shader/recompiler/decompiler/ShaderDecoder.h"
#include "graphics/shader/shaderVertexMetadata.h"
#include "libs/errno.h"
#include "spirv-tools/libspirv.h"
@@ -828,11 +828,10 @@ static void ShaderGetStaticInputInfoPS(
vs_info.stage.program != nullptr && !vs_info.stage.program->bindings.descriptors.empty()
? 1
: 0;
ps_info.push_constant_offset =
vs_info.stage.program != nullptr
? vs_info.stage.program->bindings.push_constant_offset +
vs_info.stage.program->bindings.push_constant_size
: 0;
ps_info.push_constant_offset = vs_info.stage.program != nullptr
? vs_info.stage.program->bindings.push_constant_offset +
vs_info.stage.program->bindings.push_constant_size
: 0;
for (int i = 0; i < 8; i++) {
ps_info.target_output_mode[i] = sh.target_output_mode[i];
@@ -1294,9 +1293,8 @@ static void DumpShaderRecompilerSpirv(const char* type, uint64_t shader_hash,
static std::atomic_int id = 0;
const auto base_name =
Config::GetShaderLogFolder() /
fmt::format("{:04d}_new_shader_{}_{:016x}", id++, type, shader_hash);
const auto base_name = Config::GetShaderLogFolder() /
fmt::format("{:04d}_new_shader_{}_{:016x}", id++, type, shader_hash);
Common::File::CreateDirectories(base_name.parent_path());
Common::File spv_file;
@@ -1345,9 +1343,8 @@ static void DumpShaderRecompilerOriginal(const char* type, uint64_t shader_hash,
static std::atomic_int id = 0;
const auto base_name =
Config::GetShaderLogFolder() / "original" /
fmt::format("{:04d}_new_shader_{}_{:016x}", id++, type, shader_hash);
const auto base_name = Config::GetShaderLogFolder() / "original" /
fmt::format("{:04d}_new_shader_{}_{:016x}", id++, type, shader_hash);
Common::File::CreateDirectories(base_name.parent_path());
Common::File bin_file;
+2 -2
View File
@@ -42,8 +42,8 @@ struct ShaderStageRuntime {
// Resolves an immutable native shader plan against current user data. The prior stage is preserved
// if any ReadConst, snapshot, or specialization check fails.
bool ShaderMaterializeStageRuntime(std::shared_ptr<const ShaderRecompiler::IR::Program> program,
std::span<const uint32_t> user_data, uint64_t shader_base,
ShaderStageRuntime& stage, std::string* error);
std::span<const uint32_t> user_data, uint64_t shader_base,
ShaderStageRuntime& stage, std::string* error);
struct ShaderId {
uint32_t hash0 = 0;
+2 -2
View File
@@ -6,8 +6,8 @@
namespace Libs::Graphics {
bool ShaderMaterializeStageRuntime(std::shared_ptr<const ShaderRecompiler::IR::Program> program,
std::span<const uint32_t> user_data, uint64_t shader_base,
ShaderStageRuntime& stage, std::string* error) {
std::span<const uint32_t> user_data, uint64_t shader_base,
ShaderStageRuntime& stage, std::string* error) {
if (program == nullptr) {
if (error != nullptr) {
*error = "missing native shader plan";
+1 -1
View File
@@ -17,7 +17,7 @@ bool Fail(std::string* error, const char* message) {
} // namespace
bool ShaderReadVertexMetadata(const ShaderMappedData& data, uint32_t max_user_sgprs,
ShaderVertexMetadata& metadata, std::string* error) {
ShaderVertexMetadata& metadata, std::string* error) {
if (data.user_data == nullptr) {
return Fail(error, "missing AGC user-data header");
}
+1 -1
View File
@@ -17,7 +17,7 @@ struct ShaderVertexMetadata {
// Copies the small AGC metadata subset used by the vertex path after validating every guest range.
bool ShaderReadVertexMetadata(const ShaderMappedData& data, uint32_t max_user_sgprs,
ShaderVertexMetadata& metadata, std::string* error);
ShaderVertexMetadata& metadata, std::string* error);
} // namespace Libs::Graphics
+8 -11
View File
@@ -33,8 +33,8 @@ static uint64_t MonotonicTimeNs() {
}
static std::unordered_map<KernelEqueue, KernelEqueueRef> g_equeues;
static Common::Mutex g_equeues_mutex;
static uint64_t g_next_equeue = 1;
static Common::Mutex g_equeues_mutex;
static uint64_t g_next_equeue = 1;
class KernelEqueuePrivate {
public:
@@ -459,8 +459,7 @@ int KYTY_SYSV_ABI KernelAddUserEvent(KernelEqueue eq, int id) {
int KYTY_SYSV_ABI KernelAddUserEventEdge(KernelEqueue eq, int id) {
PRINT_NAME();
LOGF("\t user event edge add: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq),
id);
LOGF("\t user event edge add: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq), id);
KernelEqueueEvent event {};
event.event.ident = static_cast<uintptr_t>(id);
@@ -485,7 +484,7 @@ int KYTY_SYSV_ABI KernelTriggerUserEvent(KernelEqueue eq, int id, void* udata) {
}
int KYTY_SYSV_ABI KernelTriggerUserEventForAll(int id, void* udata) {
int triggered = 0;
int triggered = 0;
std::vector<KernelEqueueRef> queues;
{
@@ -507,8 +506,7 @@ int KYTY_SYSV_ABI KernelTriggerUserEventForAll(int id, void* udata) {
int KYTY_SYSV_ABI KernelDeleteUserEvent(KernelEqueue eq, int id) {
PRINT_NAME();
LOGF("\t user event delete: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq),
id);
LOGF("\t user event delete: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq), id);
return KernelDeleteEvent(eq, static_cast<uintptr_t>(id), KERNEL_EVFILT_USER);
}
@@ -577,8 +575,7 @@ int KYTY_SYSV_ABI KernelAddAmprSystemEvent(KernelEqueue eq, int id, void* udata)
int KYTY_SYSV_ABI KernelDeleteAmprEvent(KernelEqueue eq, int id) {
PRINT_NAME();
LOGF("\t AMPR event delete: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq),
id);
LOGF("\t AMPR event delete: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq), id);
if (eq != KERNEL_EQUEUE_INVALID) {
(void)KernelDeleteEvent(eq, static_cast<uintptr_t>(id), KERNEL_EVFILT_USER);
@@ -590,8 +587,8 @@ int KYTY_SYSV_ABI KernelDeleteAmprEvent(KernelEqueue eq, int id) {
int KYTY_SYSV_ABI KernelDeleteAmprSystemEvent(KernelEqueue eq, int id) {
PRINT_NAME();
LOGF("\t AMPR system event delete: eq = 0x%016" PRIx64 ", id = %d\n",
static_cast<uint64_t>(eq), id);
LOGF("\t AMPR system event delete: eq = 0x%016" PRIx64 ", id = %d\n", static_cast<uint64_t>(eq),
id);
return KernelDeleteAmprEvent(eq, id);
}
+1 -1
View File
@@ -41,7 +41,7 @@ struct KernelEvent {
};
struct KernelFilter {
void* data = nullptr;
void* data = nullptr;
std::shared_ptr<void> owner;
trigger_func_t trigger_func = nullptr;
reset_func_t reset_func = nullptr;
+2 -2
View File
@@ -206,8 +206,8 @@ bool ConfigurationItem::operator<(const QTreeWidgetItem& other) const {
GetStatusText(other_item->m_info->game_status);
case GameVersionColumn:
case FirmwareVersionColumn: {
const auto& version = column == GameVersionColumn ? m_info->gameVersion
: m_info->firmwareVer;
const auto& version =
column == GameVersionColumn ? m_info->gameVersion : m_info->firmwareVer;
const auto& other_version = column == GameVersionColumn
? other_item->m_info->gameVersion
: other_item->m_info->firmwareVer;
+1 -1
View File
@@ -1,6 +1,5 @@
#include "configurationListWidget.h"
#include "patchesDialog.h"
#include "common.h"
#include "compatibilityDatabase.h"
#include "configuration.h"
@@ -8,6 +7,7 @@
#include "configurationItem.h"
#include "gameListTreeWidget.h"
#include "mainDialog.h"
#include "patchesDialog.h"
#include "trophyViewerDialog.h"
#include <QAbstractItemModel>
+16 -8
View File
@@ -272,10 +272,18 @@ static bool FindTerminal(QString* program, QStringList* prefix) {
};
static const TerminalSpec candidates[] = {
{"x-terminal-emulator", "-e"}, {"gnome-terminal", "--"}, {"konsole", "-e"},
{"xfce4-terminal", "-x"}, {"mate-terminal", "--"}, {"tilix", "-e"},
{"alacritty", "-e"}, {"kitty", nullptr}, {"foot", nullptr},
{"wezterm", "-e"}, {"urxvt", "-e"}, {"xterm", "-e"},
{"x-terminal-emulator", "-e"},
{"gnome-terminal", "--"},
{"konsole", "-e"},
{"xfce4-terminal", "-x"},
{"mate-terminal", "--"},
{"tilix", "-e"},
{"alacritty", "-e"},
{"kitty", nullptr},
{"foot", nullptr},
{"wezterm", "-e"},
{"urxvt", "-e"},
{"xterm", "-e"},
};
const auto try_candidate = [program, prefix](const QString& executable, const char* separator) {
@@ -293,7 +301,7 @@ static bool FindTerminal(QString* program, QStringList* prefix) {
if (const auto from_env = qEnvironmentVariable("TERMINAL"); !from_env.isEmpty()) {
// Reuse the known separator for an explicit terminal.
const auto env_name = QFileInfo(from_env).fileName();
const auto env_name = QFileInfo(from_env).fileName();
const char* separator = "-e";
for (const auto& candidate: candidates) {
if (env_name == QLatin1String(candidate.executable)) {
@@ -379,9 +387,9 @@ void MainDialog::RunInterpreter(QProcess* process, const Configuration& info) {
#if !defined(_WIN32)
// Report immediate launch failures.
if (!process->waitForStarted(5000)) {
QMessageBox::critical(this, tr("Error"),
tr("Failed to start:\n%1\n\n%2")
.arg(process->program(), process->errorString()));
QMessageBox::critical(
this, tr("Error"),
tr("Failed to start:\n%1\n\n%2").arg(process->program(), process->errorString()));
return;
}
#endif
+5 -7
View File
@@ -59,16 +59,14 @@ void PatchesDialog::Load() {
return;
}
const auto patches = QJsonDocument::fromJson(file.readAll())
.object()
.value(QStringLiteral("patches"))
.toArray();
const auto patches =
QJsonDocument::fromJson(file.readAll()).object().value(QStringLiteral("patches")).toArray();
for (const auto& value: patches) {
const auto patch = value.toObject();
auto* item = new QListWidgetItem(patch.value(QStringLiteral("name")).toString(), m_patches);
item->setFlags(item->flags() | Qt::ItemIsUserCheckable);
item->setCheckState(patch.value(QStringLiteral("enabled")).toBool(true) ? Qt::Checked
: Qt::Unchecked);
: Qt::Unchecked);
}
m_apply->setEnabled(!patches.isEmpty());
@@ -84,8 +82,8 @@ void PatchesDialog::Save() {
auto document = QJsonDocument::fromJson(input.readAll());
input.close();
auto root = document.object();
auto patches = root.value(QStringLiteral("patches")).toArray();
auto root = document.object();
auto patches = root.value(QStringLiteral("patches")).toArray();
for (int index = 0; index < patches.size(); index++) {
auto patch = patches[index].toObject();
patch.insert(QStringLiteral("enabled"),
+1 -1
View File
@@ -52,7 +52,7 @@ KYTY_SUBSYSTEM_INIT(Graphics) {
auto& presenter = WindowInit(width, height);
auto& video_out = VideoOut::VideoOutInit(width, height, presenter);
g_renderer = &presenter.Renderer();
g_renderer = &presenter.Renderer();
g_renderer->InitializeGpu(&video_out);
ShaderInit();
}
+2 -2
View File
@@ -1511,8 +1511,8 @@ static int ExecuteAprCommandBuffer(uint64_t command_buffer, int32_t* execution_r
} break;
case CommandKind::KernelEvent: {
const auto& command = state.kernel_event_commands[entry.index];
const auto eq = static_cast<LibKernel::EventQueue::KernelEqueue>(command.eq);
auto result = LibKernel::EventQueue::KernelTriggerUserEvent(
const auto eq = static_cast<LibKernel::EventQueue::KernelEqueue>(command.eq);
auto result = LibKernel::EventQueue::KernelTriggerUserEvent(
eq, command.id, reinterpret_cast<void*>(command.data));
if (result != OK) {
LOGF("\tAPR submit event failed: eq=0x%016" PRIx64 ", id=%" PRId32
+66 -4
View File
@@ -40,7 +40,10 @@
#define NOMINMAX
#endif
#include <windows.h>
#elif !defined(__APPLE__)
#elif defined(__APPLE__)
#include <csignal>
#include <sys/ucontext.h>
#else
#include <csignal>
#include <ucontext.h>
#endif
@@ -837,7 +840,7 @@ static void ApplySignalUcontext(CONTEXT* dst_ctx, const SignalUcontext& src_ctx)
}
#endif
#if KYTY_PLATFORM != KYTY_PLATFORM_WINDOWS && !defined(__APPLE__) && defined(__x86_64__)
#if KYTY_PLATFORM != KYTY_PLATFORM_WINDOWS && defined(__x86_64__)
static SignalUcontext CreateSignalUcontextFromHost(const ucontext_t* host_ctx) {
SignalUcontext ctx = {};
@@ -845,6 +848,34 @@ static SignalUcontext CreateSignalUcontextFromHost(const ucontext_t* host_ctx) {
return ctx;
}
#if defined(__APPLE__)
const auto& ss = host_ctx->uc_mcontext->__ss;
ctx.uc_mcontext.mc_rdi = ss.__rdi;
ctx.uc_mcontext.mc_rsi = ss.__rsi;
ctx.uc_mcontext.mc_rdx = ss.__rdx;
ctx.uc_mcontext.mc_rcx = ss.__rcx;
ctx.uc_mcontext.mc_r8 = ss.__r8;
ctx.uc_mcontext.mc_r9 = ss.__r9;
ctx.uc_mcontext.mc_rax = ss.__rax;
ctx.uc_mcontext.mc_rbx = ss.__rbx;
ctx.uc_mcontext.mc_rbp = ss.__rbp;
ctx.uc_mcontext.mc_r10 = ss.__r10;
ctx.uc_mcontext.mc_r11 = ss.__r11;
ctx.uc_mcontext.mc_r12 = ss.__r12;
ctx.uc_mcontext.mc_r13 = ss.__r13;
ctx.uc_mcontext.mc_r14 = ss.__r14;
ctx.uc_mcontext.mc_r15 = ss.__r15;
ctx.uc_mcontext.mc_rip = ss.__rip;
ctx.uc_mcontext.mc_rsp = ss.__rsp;
ctx.uc_mcontext.mc_rflags = ss.__rflags;
ctx.uc_mcontext.mc_cs = ss.__cs & 0xffffu;
ctx.uc_mcontext.mc_gs = static_cast<uint16_t>(ss.__gs & 0xffffu);
ctx.uc_mcontext.mc_fs = static_cast<uint16_t>(ss.__fs & 0xffffu);
ctx.uc_mcontext.mc_len = sizeof(SignalMcontext);
return ctx;
#else
const auto* gregs = host_ctx->uc_mcontext.gregs;
ctx.uc_mcontext.mc_rdi = static_cast<uint64_t>(gregs[REG_RDI]);
@@ -874,6 +905,7 @@ static SignalUcontext CreateSignalUcontextFromHost(const ucontext_t* host_ctx) {
ctx.uc_mcontext.mc_len = sizeof(SignalMcontext);
return ctx;
#endif
}
static void ApplySignalUcontextToHost(ucontext_t* dst_ctx, const SignalUcontext& src_ctx) {
@@ -881,6 +913,29 @@ static void ApplySignalUcontextToHost(ucontext_t* dst_ctx, const SignalUcontext&
return;
}
#if defined(__APPLE__)
auto& ss = dst_ctx->uc_mcontext->__ss;
ss.__rdi = src_ctx.uc_mcontext.mc_rdi;
ss.__rsi = src_ctx.uc_mcontext.mc_rsi;
ss.__rdx = src_ctx.uc_mcontext.mc_rdx;
ss.__rcx = src_ctx.uc_mcontext.mc_rcx;
ss.__r8 = src_ctx.uc_mcontext.mc_r8;
ss.__r9 = src_ctx.uc_mcontext.mc_r9;
ss.__rax = src_ctx.uc_mcontext.mc_rax;
ss.__rbx = src_ctx.uc_mcontext.mc_rbx;
ss.__rbp = src_ctx.uc_mcontext.mc_rbp;
ss.__r10 = src_ctx.uc_mcontext.mc_r10;
ss.__r11 = src_ctx.uc_mcontext.mc_r11;
ss.__r12 = src_ctx.uc_mcontext.mc_r12;
ss.__r13 = src_ctx.uc_mcontext.mc_r13;
ss.__r14 = src_ctx.uc_mcontext.mc_r14;
ss.__r15 = src_ctx.uc_mcontext.mc_r15;
ss.__rip = src_ctx.uc_mcontext.mc_rip;
ss.__rsp = src_ctx.uc_mcontext.mc_rsp;
ss.__rflags = src_ctx.uc_mcontext.mc_rflags;
// Segment selectors are left untouched; XNU validates them on sigreturn.
#else
auto* gregs = dst_ctx->uc_mcontext.gregs;
gregs[REG_RDI] = static_cast<greg_t>(src_ctx.uc_mcontext.mc_rdi);
@@ -903,10 +958,17 @@ static void ApplySignalUcontextToHost(ucontext_t* dst_ctx, const SignalUcontext&
gregs[REG_EFL] = static_cast<greg_t>(src_ctx.uc_mcontext.mc_rflags);
// The kernel validates packed segment selectors on sigreturn.
#endif
}
static int SignalDispatchHostSignal() {
#if defined(__APPLE__)
// macOS has no realtime signals; SIGUSR1 is otherwise unused on the host side (the
// guest's SIGUSR1 is an emulated signal number, not a host registration).
static const int host_signal = SIGUSR1;
#else
static const int host_signal = SIGRTMIN + 3;
#endif
return host_signal;
}
@@ -1174,12 +1236,12 @@ static int KYTY_SYSV_ABI KernelRaiseException(Pthread thread, int signum) {
}
CloseHandle(target_thread);
return OK;
#elif !defined(__APPLE__) && defined(__x86_64__)
#elif defined(__x86_64__)
// Deliver on the target thread.
if (thread == PthreadSelfOrNull()) {
SignalDispatchScope scope;
auto ctx = CreateCurrentGuestCallSignalUcontext(
reinterpret_cast<uint64_t>(__builtin_return_address(0)));
reinterpret_cast<uint64_t>(__builtin_return_address(0)));
handler(signum, &ctx);
return OK;
}
+10 -8
View File
@@ -438,11 +438,12 @@ int KYTY_SYSV_ABI SaveDataMount3(const SaveDataMount3* mount, SaveDataMountResul
*mount_result = {};
Common::LockGuard lock(g_mount_mutex);
const std::string dir_name = mount->dir_name->data;
const std::string mount_dir = std::string(SAVE_DATA_DIR) + "/" + get_title_id() + "/" + dir_name;
const bool create = ((mount->mount_mode & 4u) != 0);
const bool create2 = ((mount->mount_mode & 32u) != 0);
const bool open = (!create && !create2 && ((mount->mount_mode & 3u) != 0));
const std::string dir_name = mount->dir_name->data;
const std::string mount_dir =
std::string(SAVE_DATA_DIR) + "/" + get_title_id() + "/" + dir_name;
const bool create = ((mount->mount_mode & 4u) != 0);
const bool create2 = ((mount->mount_mode & 32u) != 0);
const bool open = (!create && !create2 && ((mount->mount_mode & 3u) != 0));
const int slot = g_mount_slots.FindAvailable(dir_name);
if (slot == SaveDataMountSlots::BUSY) {
@@ -594,9 +595,10 @@ int KYTY_SYSV_ABI SaveDataTransferringMount(const SaveDataTransferringMount* mou
*mount_result = {};
Common::LockGuard lock(g_mount_mutex);
const std::string dir_name = mount->dir_name->data;
const std::string mount_dir = std::string(SAVE_DATA_DIR) + "/" + get_title_id() + "/" + dir_name;
const int slot = g_mount_slots.FindAvailable(dir_name);
const std::string dir_name = mount->dir_name->data;
const std::string mount_dir =
std::string(SAVE_DATA_DIR) + "/" + get_title_id() + "/" + dir_name;
const int slot = g_mount_slots.FindAvailable(dir_name);
if (slot == SaveDataMountSlots::BUSY) {
return SAVE_DATA_ERROR_BUSY;
}
+31
View File
@@ -0,0 +1,31 @@
#include "common/abi.h"
#include "libs/errno.h"
#include "libs/libs.h"
#include "loader/symbolDatabase.h"
namespace Libs {
LIB_VERSION("TextToSpeech2", 1, "TextToSpeech2", 1, 1);
namespace TextToSpeech2 {
static int KYTY_SYSV_ABI TextToSpeech2GetSpeechStatus() {
PRINT_NAME();
return OK;
}
static int KYTY_SYSV_ABI TextToSpeech2Cancel() {
PRINT_NAME();
return OK;
}
} // namespace TextToSpeech2
LIB_DEFINE(InitTextToSpeech2_1) {
LIB_FUNC("08JSg9p6bgQ", TextToSpeech2::TextToSpeech2GetSpeechStatus);
LIB_FUNC("2jiIxUmcsGo", TextToSpeech2::TextToSpeech2Cancel);
}
} // namespace Libs
+252 -36
View File
@@ -1,10 +1,12 @@
#include "common/abi.h"
#include "libs/errno.h"
#include "libs/libs.h"
#include "libs/videoDec2Decoder.h"
#include "loader/symbolDatabase.h"
#include <cstddef>
#include <cstdint>
#include <cstring>
#include <mutex>
#include <unordered_set>
@@ -14,6 +16,7 @@ LIB_VERSION("Videodec2", 1, "Videodec2", 1, 1);
namespace VideoDec2 {
constexpr int32_t VIDEODEC2_ERROR_API_FAIL = -2128805632; // 0x811d0100
constexpr int32_t VIDEODEC2_ERROR_STRUCT_SIZE = -2128805631; // 0x811d0101
constexpr int32_t VIDEODEC2_ERROR_ARGUMENT_POINTER = -2128805630; // 0x811d0102
constexpr int32_t VIDEODEC2_ERROR_DECODER_INSTANCE = -2128805629; // 0x811d0103
@@ -21,13 +24,20 @@ constexpr int32_t VIDEODEC2_ERROR_MEMORY_SIZE = -2128805628; // 0x811d0
constexpr int32_t VIDEODEC2_ERROR_MEMORY_POINTER = -2128805627; // 0x811d0105
constexpr int32_t VIDEODEC2_ERROR_FRAME_BUFFER_SIZE = -2128805626; // 0x811d0106
constexpr int32_t VIDEODEC2_ERROR_FRAME_BUFFER_POINTER = -2128805625; // 0x811d0107
constexpr int32_t VIDEODEC2_ERROR_ACCESS_UNIT_SIZE = -2128805619; // 0x811d010d
constexpr int32_t VIDEODEC2_ERROR_ACCESS_UNIT_POINTER = -2128805618; // 0x811d010e
constexpr int32_t VIDEODEC2_ERROR_OUTPUT_INFO = -2128805617; // 0x811d010f
constexpr int32_t VIDEODEC2_ERROR_COMPUTE_QUEUE = -2128805616; // 0x811d0110
constexpr int32_t VIDEODEC2_ERROR_CONFIG_INFO = -2128805376; // 0x811d0200
constexpr int32_t VIDEODEC2_ERROR_COMPUTE_PIPE_ID = -2128805375; // 0x811d0201
constexpr int32_t VIDEODEC2_ERROR_COMPUTE_QUEUE_ID = -2128805374; // 0x811d0202
constexpr int32_t VIDEODEC2_ERROR_RESOURCE_TYPE = -2128805373; // 0x811d0203
constexpr int32_t VIDEODEC2_ERROR_CODEC_TYPE = -2128805372; // 0x811d0204
constexpr int32_t VIDEODEC2_ERROR_INPUT_QUEUE_DEPTH = -2128805370; // 0x811d0206
constexpr int32_t VIDEODEC2_ERROR_DPB_FRAME_COUNT = -2128805367; // 0x811d0209
constexpr int32_t VIDEODEC2_ERROR_FRAME_WIDTH_HEIGHT = -2128805366; // 0x811d020a
constexpr int32_t VIDEODEC2_ERROR_ACCESS_UNIT = -2128805119; // 0x811d0301
constexpr int32_t VIDEODEC2_ERROR_OVERSIZE_DECODE = -2128805118; // 0x811d0302
constexpr uint32_t VIDEODEC2_RESOURCE_TYPE_COMPUTE = 1;
constexpr size_t VIDEODEC2_MIN_MEMORY_SIZE = 16ull * 1024ull * 1024ull;
@@ -101,6 +111,70 @@ struct Videodec2FrameBuffer {
bool is_accepted;
};
struct Videodec2AvcPictureInfo {
size_t this_size;
bool is_valid;
uint64_t pts_data;
uint64_t dts_data;
uint64_t attached_data;
uint8_t idr_picture_flag;
uint8_t profile_idc;
uint8_t level_idc;
uint32_t pic_width_in_mbs_minus1;
uint32_t pic_height_in_map_units_minus1;
uint8_t frame_mbs_only_flag;
uint8_t frame_cropping_flag;
uint32_t frame_crop_left_offset;
uint32_t frame_crop_right_offset;
uint32_t frame_crop_top_offset;
uint32_t frame_crop_bottom_offset;
uint8_t aspect_ratio_info_present_flag;
uint8_t aspect_ratio_idc;
uint16_t sar_width;
uint16_t sar_height;
uint8_t video_signal_type_present_flag;
uint8_t video_format;
uint8_t video_full_range_flag;
uint8_t colour_description_present_flag;
uint8_t colour_primaries;
uint8_t transfer_characteristics;
uint8_t matrix_coefficients;
uint8_t timing_info_present_flag;
uint32_t num_units_in_tick;
uint32_t time_scale;
uint8_t fixed_frame_rate_flag;
uint8_t bitstream_restriction_flag;
uint8_t max_dec_frame_buffering;
uint8_t pic_struct_present_flag;
uint8_t pic_struct;
uint8_t field_pic_flag;
uint8_t bottom_field_flag;
uint8_t sequence_parameter_set_present_flag;
uint8_t picture_parameter_set_present_flag;
uint8_t au_delimiter_present_flag;
uint8_t end_of_sequence_present_flag;
uint8_t end_of_stream_present_flag;
uint8_t filler_data_present_flag;
uint8_t picture_timing_sei_present_flag;
uint8_t buffering_period_sei_present_flag;
uint8_t constraint_set0_flag;
uint8_t constraint_set1_flag;
uint8_t constraint_set2_flag;
uint8_t constraint_set3_flag;
uint8_t constraint_set4_flag;
uint8_t constraint_set5_flag;
};
struct Videodec2ComputeMemoryInfo {
size_t this_size;
size_t cpu_gpu_memory_size;
@@ -116,10 +190,7 @@ struct Videodec2ComputeConfigInfo {
uint16_t reserved1;
};
struct DecoderState {
uint64_t magic;
uint32_t codec_type;
};
using DecoderState = Decoder::Instance;
static_assert(sizeof(Videodec2ComputeMemoryInfo) == 24);
static_assert(sizeof(Videodec2ComputeConfigInfo) == 16);
@@ -128,8 +199,7 @@ static_assert(sizeof(Videodec2DecoderMemoryInfo) == 72);
static_assert(sizeof(Videodec2InputData) == 48);
static_assert(sizeof(Videodec2OutputInfo) == 56);
static_assert(sizeof(Videodec2FrameBuffer) == 32);
constexpr uint64_t DECODER_MAGIC = 0x4b59545956444543ull; // KYTYVDEC
static_assert(sizeof(Videodec2AvcPictureInfo) == 120);
static std::mutex g_decoder_mutex;
static std::unordered_set<void*> g_decoders;
@@ -156,15 +226,54 @@ static void FillNoPictureOutput(const Videodec2FrameBuffer* frame_buffer,
output_info->frame_height = 0;
output_info->frame_buffer = frame_buffer != nullptr ? frame_buffer->frame_buffer : nullptr;
output_info->frame_buffer_size = frame_buffer != nullptr ? frame_buffer->frame_buffer_size : 0;
output_info->frame_format = VIDEODEC2_FRAME_FORMAT_DEFAULT;
output_info->frame_pitch_in_bytes = 0;
if (output_info->this_size == sizeof(Videodec2OutputInfo)) {
output_info->frame_format = VIDEODEC2_FRAME_FORMAT_DEFAULT;
output_info->frame_pitch_in_bytes = 0;
}
}
static int32_t ValidateDecoderConfig(const Videodec2DecoderConfigInfo* config) {
static int32_t MapDecoderResult(Decoder::Result result) {
switch (result) {
case Decoder::Result::Ok: return OK;
case Decoder::Result::ApiFail: return VIDEODEC2_ERROR_API_FAIL;
case Decoder::Result::AccessUnit: return VIDEODEC2_ERROR_ACCESS_UNIT;
case Decoder::Result::FrameBufferSize: return VIDEODEC2_ERROR_FRAME_BUFFER_SIZE;
case Decoder::Result::OversizeDecode: return VIDEODEC2_ERROR_OVERSIZE_DECODE;
}
return VIDEODEC2_ERROR_API_FAIL;
}
static void ApplyDecodedOutput(const Decoder::Output& decoded, Videodec2FrameBuffer* frame_buffer,
Videodec2OutputInfo* output_info) {
frame_buffer->is_accepted = decoded.buffer_accepted;
if (!decoded.valid) {
return;
}
output_info->is_valid = true;
output_info->is_error_frame = decoded.error_frame;
output_info->picture_count = 1;
output_info->codec_type = decoded.codec_type;
output_info->frame_width = decoded.width;
output_info->frame_pitch = decoded.pitch;
output_info->frame_height = decoded.height;
output_info->frame_buffer = decoded.buffer;
output_info->frame_buffer_size = decoded.buffer_size;
if (output_info->this_size == sizeof(Videodec2OutputInfo)) {
output_info->frame_format = VIDEODEC2_FRAME_FORMAT_DEFAULT;
output_info->frame_pitch_in_bytes = decoded.pitch;
}
}
static int32_t ValidateDecoderConfig(const Videodec2DecoderConfigInfo* config,
bool require_compute_queue) {
if (config->resource_type != VIDEODEC2_RESOURCE_TYPE_COMPUTE) {
return VIDEODEC2_ERROR_RESOURCE_TYPE;
}
if (!Decoder::IsCodecSupported(config->codec_type)) {
return VIDEODEC2_ERROR_CODEC_TYPE;
}
if (config->reserved0 != 0 || config->reserved1 != 0) {
return VIDEODEC2_ERROR_CONFIG_INFO;
}
@@ -182,8 +291,8 @@ static int32_t ValidateDecoderConfig(const Videodec2DecoderConfigInfo* config) {
return VIDEODEC2_ERROR_FRAME_WIDTH_HEIGHT;
}
if (config->compute_queue == nullptr) {
return VIDEODEC2_ERROR_CONFIG_INFO;
if (require_compute_queue && config->compute_queue == nullptr) {
return VIDEODEC2_ERROR_COMPUTE_QUEUE;
}
return OK;
@@ -243,7 +352,6 @@ static int32_t KYTY_SYSV_ABI AllocateComputeQueue(
}
*compute_queue = compute_memory_info->cpu_gpu_memory;
return OK;
}
@@ -266,7 +374,7 @@ static int32_t KYTY_SYSV_ABI QueryDecoderMemoryInfo(const Videodec2DecoderConfig
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
const auto validation_result = ValidateDecoderConfig(config);
const auto validation_result = ValidateDecoderConfig(config, false);
if (validation_result != OK) {
return validation_result;
}
@@ -298,7 +406,7 @@ static int32_t KYTY_SYSV_ABI CreateDecoder(const Videodec2DecoderConfigInfo* con
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
const auto validation_result = ValidateDecoderConfig(config);
const auto validation_result = ValidateDecoderConfig(config, true);
if (validation_result != OK) {
return validation_result;
}
@@ -315,9 +423,11 @@ static int32_t KYTY_SYSV_ABI CreateDecoder(const Videodec2DecoderConfigInfo* con
return VIDEODEC2_ERROR_MEMORY_POINTER;
}
auto* state = new DecoderState {};
state->magic = DECODER_MAGIC;
state->codec_type = config->codec_type;
auto* state =
Decoder::Create({config->codec_type, config->max_frame_width, config->max_frame_height});
if (state == nullptr) {
return VIDEODEC2_ERROR_API_FAIL;
}
{
std::scoped_lock lock(g_decoder_mutex);
@@ -325,7 +435,6 @@ static int32_t KYTY_SYSV_ABI CreateDecoder(const Videodec2DecoderConfigInfo* con
}
*decoder = state;
return OK;
}
@@ -343,7 +452,7 @@ static int32_t KYTY_SYSV_ABI DeleteDecoder(Videodec2Decoder decoder) {
g_decoders.erase(it);
}
delete state;
Decoder::Destroy(state);
return OK;
}
@@ -353,8 +462,8 @@ static int32_t KYTY_SYSV_ABI Decode(Videodec2Decoder decoder, const Videodec2Inp
Videodec2OutputInfo* output_info) {
PRINT_NAME();
const auto* state = GetDecoder(decoder);
if (state == nullptr || state->magic != DECODER_MAGIC) {
auto* state = GetDecoder(decoder);
if (state == nullptr) {
return VIDEODEC2_ERROR_DECODER_INSTANCE;
}
@@ -368,8 +477,12 @@ static int32_t KYTY_SYSV_ABI Decode(Videodec2Decoder decoder, const Videodec2Inp
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
if (input_data->au_size != 0 && input_data->au_data == nullptr) {
return VIDEODEC2_ERROR_ARGUMENT_POINTER;
if (input_data->au_size == 0) {
return VIDEODEC2_ERROR_ACCESS_UNIT_SIZE;
}
if (input_data->au_data == nullptr) {
return VIDEODEC2_ERROR_ACCESS_UNIT_POINTER;
}
if (frame_buffer->frame_buffer_size == 0) {
@@ -381,17 +494,24 @@ static int32_t KYTY_SYSV_ABI Decode(Videodec2Decoder decoder, const Videodec2Inp
}
frame_buffer->is_accepted = false;
FillNoPictureOutput(frame_buffer, output_info, state->codec_type);
FillNoPictureOutput(frame_buffer, output_info, Decoder::GetCodecType(state));
return OK;
Decoder::Output decoded {};
const auto result =
Decoder::Decode(state,
{input_data->au_data, input_data->au_size, input_data->pts_data,
input_data->dts_data, input_data->attached_data},
{frame_buffer->frame_buffer, frame_buffer->frame_buffer_size}, &decoded);
ApplyDecodedOutput(decoded, frame_buffer, output_info);
return MapDecoderResult(result);
}
static int32_t KYTY_SYSV_ABI Flush(Videodec2Decoder decoder, Videodec2FrameBuffer* frame_buffer,
Videodec2OutputInfo* output_info) {
PRINT_NAME();
const auto* state = GetDecoder(decoder);
if (state == nullptr || state->magic != DECODER_MAGIC) {
auto* state = GetDecoder(decoder);
if (state == nullptr) {
return VIDEODEC2_ERROR_DECODER_INSTANCE;
}
@@ -404,26 +524,40 @@ static int32_t KYTY_SYSV_ABI Flush(Videodec2Decoder decoder, Videodec2FrameBuffe
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
frame_buffer->is_accepted = false;
FillNoPictureOutput(frame_buffer, output_info, state->codec_type);
if (frame_buffer->frame_buffer_size == 0) {
return VIDEODEC2_ERROR_FRAME_BUFFER_SIZE;
}
return OK;
if (frame_buffer->frame_buffer == nullptr) {
return VIDEODEC2_ERROR_FRAME_BUFFER_POINTER;
}
frame_buffer->is_accepted = false;
FillNoPictureOutput(frame_buffer, output_info, Decoder::GetCodecType(state));
Decoder::Output decoded {};
const auto result = Decoder::Flush(
state, {frame_buffer->frame_buffer, frame_buffer->frame_buffer_size}, &decoded);
ApplyDecodedOutput(decoded, frame_buffer, output_info);
return MapDecoderResult(result);
}
static int32_t KYTY_SYSV_ABI Reset(Videodec2Decoder decoder) {
PRINT_NAME();
const auto* state = GetDecoder(decoder);
return state != nullptr && state->magic == DECODER_MAGIC ? OK
: VIDEODEC2_ERROR_DECODER_INSTANCE;
auto* state = GetDecoder(decoder);
if (state == nullptr) {
return VIDEODEC2_ERROR_DECODER_INSTANCE;
}
Decoder::Reset(state);
return OK;
}
static int32_t KYTY_SYSV_ABI GetPictureInfo(const Videodec2OutputInfo* output_info,
void* /*first_picture_info*/,
void* /*second_picture_info*/) {
void* first_picture_info, void* second_picture_info) {
PRINT_NAME();
if (output_info == nullptr) {
if (output_info == nullptr || first_picture_info == nullptr) {
return VIDEODEC2_ERROR_ARGUMENT_POINTER;
}
@@ -431,6 +565,88 @@ static int32_t KYTY_SYSV_ABI GetPictureInfo(const Videodec2OutputInfo* output_in
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
if (!output_info->is_valid || output_info->picture_count == 0 ||
output_info->frame_buffer == nullptr) {
return VIDEODEC2_ERROR_OUTPUT_INFO;
}
Decoder::PictureInfo decoded {};
if (!Decoder::GetPictureInfo(output_info->frame_buffer, &decoded) ||
decoded.codec_type != output_info->codec_type) {
return VIDEODEC2_ERROR_OUTPUT_INFO;
}
auto fill_common = [&decoded](void* destination, bool valid) -> int32_t {
auto* bytes = static_cast<uint8_t*>(destination);
const auto size = *static_cast<const size_t*>(destination);
if (size < 40 || size > 256) {
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
std::memset(bytes + sizeof(size_t), 0, size - sizeof(size_t));
bytes[8] = valid ? 1 : 0;
if (valid) {
std::memcpy(bytes + 16, &decoded.pts, sizeof(decoded.pts));
std::memcpy(bytes + 24, &decoded.dts, sizeof(decoded.dts));
std::memcpy(bytes + 32, &decoded.attached_data, sizeof(decoded.attached_data));
}
return OK;
};
if (output_info->codec_type == 1) {
const auto requested_size = *static_cast<const size_t*>(first_picture_info);
if (requested_size != sizeof(Videodec2AvcPictureInfo) &&
(requested_size | 16u) != sizeof(Videodec2AvcPictureInfo)) {
return VIDEODEC2_ERROR_STRUCT_SIZE;
}
Videodec2AvcPictureInfo picture {};
picture.this_size = requested_size;
picture.is_valid = true;
picture.pts_data = decoded.pts;
picture.dts_data = decoded.dts;
picture.attached_data = decoded.attached_data;
picture.idr_picture_flag = decoded.key_frame ? 1 : 0;
picture.profile_idc = static_cast<uint8_t>(decoded.profile);
picture.level_idc = static_cast<uint8_t>(decoded.level);
picture.pic_width_in_mbs_minus1 = (decoded.width + 15u) / 16u - 1u;
picture.pic_height_in_map_units_minus1 = (decoded.height + 15u) / 16u - 1u;
picture.frame_mbs_only_flag = 1;
picture.frame_cropping_flag = decoded.crop_left != 0 || decoded.crop_right != 0 ||
decoded.crop_top != 0 || decoded.crop_bottom != 0
? 1
: 0;
picture.frame_crop_left_offset = decoded.crop_left;
picture.frame_crop_right_offset = decoded.crop_right;
picture.frame_crop_top_offset = decoded.crop_top;
picture.frame_crop_bottom_offset = decoded.crop_bottom;
picture.aspect_ratio_info_present_flag =
decoded.sar_width != 0 && decoded.sar_height != 0 ? 1 : 0;
picture.aspect_ratio_idc = picture.aspect_ratio_info_present_flag ? 255 : 0;
picture.sar_width = decoded.sar_width;
picture.sar_height = decoded.sar_height;
picture.video_signal_type_present_flag = 1;
picture.video_format = 5;
picture.video_full_range_flag = decoded.color_range == 2 ? 1 : 0;
picture.colour_description_present_flag =
decoded.color_primaries != 0 || decoded.color_trc != 0 || decoded.color_space != 0 ? 1
: 0;
picture.colour_primaries = decoded.color_primaries;
picture.transfer_characteristics = decoded.color_trc;
picture.matrix_coefficients = decoded.color_space;
std::memcpy(first_picture_info, &picture, requested_size);
} else {
const auto result = fill_common(first_picture_info, true);
if (result != OK) {
return result;
}
}
if (second_picture_info != nullptr) {
const auto result = fill_common(second_picture_info, false);
if (result != OK) {
return result;
}
}
return OK;
}
+2
View File
@@ -66,6 +66,7 @@ LIB_DEFINE(InitSaveData_1);
LIB_DEFINE(InitShare_1);
LIB_DEFINE(InitSysmodule_1);
LIB_DEFINE(InitSystemService_1);
LIB_DEFINE(InitTextToSpeech2_1);
LIB_DEFINE(InitUserService_1);
LIB_DEFINE(InitVideoOut_1);
@@ -100,6 +101,7 @@ void InitAll(Loader::SymbolDatabase* s) {
LIB_LOAD(InitShare_1);
LIB_LOAD(InitSysmodule_1);
LIB_LOAD(InitSystemService_1);
LIB_LOAD(InitTextToSpeech2_1);
LIB_LOAD(LibUlt::InitUlt_1);
LIB_LOAD(InitUserService_1);
LIB_LOAD(VideoDec2::InitVideoDec2_1);
+412
View File
@@ -0,0 +1,412 @@
#include "libs/videoDec2Decoder.h"
#include "common/logging/log.h"
#include <algorithm>
#include <cstddef>
#include <cstdint>
#include <cstring>
#include <limits>
#include <mutex>
#include <unordered_map>
#include <unordered_set>
extern "C" {
#include <libavcodec/avcodec.h>
#include <libavutil/buffer.h>
#include <libavutil/error.h>
#include <libavutil/frame.h>
#include <libavutil/pixfmt.h>
#include <libswscale/swscale.h>
}
namespace Libs::VideoDec2::Decoder {
namespace {
constexpr uint32_t CODEC_TYPE_AVC = 1;
constexpr uint32_t CODEC_TYPE_HEVC = 974921;
constexpr uint32_t CODEC_TYPE_VP9 = 2382845;
struct PacketMetadata {
uint64_t pts = TIMESTAMP_INVALID;
uint64_t dts = TIMESTAMP_INVALID;
uint64_t attached_data = 0;
};
struct StoredPicture {
const Instance* owner = nullptr;
PictureInfo info;
};
std::mutex g_picture_mutex;
std::unordered_map<void*, StoredPicture> g_picture_infos;
AVCodecID GetAvCodecId(uint32_t codec_type) {
switch (codec_type) {
case CODEC_TYPE_AVC: return AV_CODEC_ID_H264;
case CODEC_TYPE_HEVC: return AV_CODEC_ID_HEVC;
case CODEC_TYPE_VP9: return AV_CODEC_ID_VP9;
default: return AV_CODEC_ID_NONE;
}
}
const char* AvErrorString(int error) {
thread_local char text[AV_ERROR_MAX_STRING_SIZE] {};
if (av_strerror(error, text, sizeof(text)) != 0) {
std::strcpy(text, "unknown FFmpeg error");
}
return text;
}
uint32_t AlignUp(uint32_t value, uint32_t alignment) {
return (value + alignment - 1u) & ~(alignment - 1u);
}
int64_t ToAvTimestamp(uint64_t timestamp) {
return timestamp == TIMESTAMP_INVALID ||
timestamp > static_cast<uint64_t>(std::numeric_limits<int64_t>::max())
? AV_NOPTS_VALUE
: static_cast<int64_t>(timestamp);
}
} // namespace
class Instance {
public:
explicit Instance(const Config& config): m_config(config) {}
~Instance() {
ClearPictureMetadata();
if (m_sws != nullptr) {
sws_freeContext(m_sws);
}
if (m_codec != nullptr) {
avcodec_free_context(&m_codec);
}
}
Instance(const Instance&) = delete;
Instance& operator=(const Instance&) = delete;
[[nodiscard]] bool Initialize() {
const AVCodec* decoder = avcodec_find_decoder(GetAvCodecId(m_config.codec_type));
if (decoder == nullptr) {
LOGF("Videodec2: FFmpeg decoder is unavailable for codec type %u\n",
m_config.codec_type);
return false;
}
m_codec = avcodec_alloc_context3(decoder);
if (m_codec == nullptr) {
LOGF("Videodec2: avcodec_alloc_context3 failed\n");
return false;
}
// This carries PTS/DTS/attachedData through codecs that reorder B frames.
m_codec->flags |= AV_CODEC_FLAG_COPY_OPAQUE;
const int result = avcodec_open2(m_codec, decoder, nullptr);
if (result < 0) {
LOGF("Videodec2: avcodec_open2 failed: %s (%d)\n", AvErrorString(result), result);
return false;
}
return true;
}
[[nodiscard]] uint32_t CodecType() const { return m_config.codec_type; }
[[nodiscard]] Result DecodeInput(const Input& input, const FrameBuffer& frame_buffer,
Output* output) {
std::scoped_lock lock(m_mutex);
*output = {};
m_draining = false;
AVPacket* packet = av_packet_alloc();
AVFrame* frame = av_frame_alloc();
if (packet == nullptr || frame == nullptr ||
input.size > static_cast<size_t>(std::numeric_limits<int>::max())) {
av_packet_free(&packet);
av_frame_free(&frame);
return Result::ApiFail;
}
int result = av_new_packet(packet, static_cast<int>(input.size));
if (result < 0) {
LOGF("Videodec2: av_new_packet failed: %s (%d)\n", AvErrorString(result), result);
av_packet_free(&packet);
av_frame_free(&frame);
return Result::ApiFail;
}
std::memcpy(packet->data, input.data, input.size);
packet->pts = ToAvTimestamp(input.pts);
packet->dts = ToAvTimestamp(input.dts);
packet->opaque_ref = av_buffer_alloc(sizeof(PacketMetadata));
if (packet->opaque_ref == nullptr) {
av_packet_free(&packet);
av_frame_free(&frame);
return Result::ApiFail;
}
const PacketMetadata metadata {input.pts, input.dts, input.attached_data};
std::memcpy(packet->opaque_ref->data, &metadata, sizeof(metadata));
bool have_pending_frame = false;
result = avcodec_send_packet(m_codec, packet);
if (result == AVERROR(EAGAIN)) {
result = avcodec_receive_frame(m_codec, frame);
if (result < 0) {
LOGF("Videodec2: decoder rejected an AU while no output was available: %s (%d)\n",
AvErrorString(result), result);
av_packet_free(&packet);
av_frame_free(&frame);
return Result::AccessUnit;
}
have_pending_frame = true;
result = avcodec_send_packet(m_codec, packet);
}
if (result < 0) {
LOGF("Videodec2: avcodec_send_packet failed: %s (%d)\n", AvErrorString(result), result);
av_packet_free(&packet);
av_frame_free(&frame);
return Result::AccessUnit;
}
Result decode_result = Result::Ok;
if (!have_pending_frame) {
result = avcodec_receive_frame(m_codec, frame);
if (result != AVERROR(EAGAIN) && result != AVERROR_EOF) {
if (result < 0) {
LOGF("Videodec2: avcodec_receive_frame failed: %s (%d)\n",
AvErrorString(result), result);
decode_result = Result::AccessUnit;
} else {
decode_result = CopyFrame(frame, frame_buffer, output);
}
}
} else {
decode_result = CopyFrame(frame, frame_buffer, output);
}
av_packet_free(&packet);
av_frame_free(&frame);
return decode_result;
}
[[nodiscard]] Result FlushOutput(const FrameBuffer& frame_buffer, Output* output) {
std::scoped_lock lock(m_mutex);
*output = {};
AVFrame* frame = av_frame_alloc();
if (frame == nullptr) {
return Result::ApiFail;
}
if (!m_draining) {
const int send_result = avcodec_send_packet(m_codec, nullptr);
if (send_result == 0 || send_result == AVERROR_EOF) {
m_draining = true;
} else if (send_result != AVERROR(EAGAIN)) {
LOGF("Videodec2: flushing decoder failed: %s (%d)\n", AvErrorString(send_result),
send_result);
av_frame_free(&frame);
return Result::ApiFail;
}
}
const int receive_result = avcodec_receive_frame(m_codec, frame);
if (receive_result == AVERROR(EAGAIN) || receive_result == AVERROR_EOF) {
av_frame_free(&frame);
return Result::Ok;
}
if (receive_result < 0) {
LOGF("Videodec2: receiving a flushed frame failed: %s (%d)\n",
AvErrorString(receive_result), receive_result);
av_frame_free(&frame);
return Result::ApiFail;
}
const auto result = CopyFrame(frame, frame_buffer, output);
av_frame_free(&frame);
return result;
}
void ResetDecoder() {
std::scoped_lock lock(m_mutex);
avcodec_flush_buffers(m_codec);
m_draining = false;
ClearPictureMetadata();
}
private:
[[nodiscard]] PictureInfo MakePictureInfo(const AVFrame* frame) const {
PictureInfo result {};
if (frame->opaque_ref != nullptr && frame->opaque_ref->size >= sizeof(PacketMetadata)) {
PacketMetadata metadata {};
std::memcpy(&metadata, frame->opaque_ref->data, sizeof(metadata));
result.pts = metadata.pts;
result.dts = metadata.dts;
result.attached_data = metadata.attached_data;
} else {
result.pts = frame->pts == AV_NOPTS_VALUE ? TIMESTAMP_INVALID
: static_cast<uint64_t>(frame->pts);
result.dts = frame->pkt_dts == AV_NOPTS_VALUE ? TIMESTAMP_INVALID
: static_cast<uint64_t>(frame->pkt_dts);
}
result.codec_type = m_config.codec_type;
result.width = static_cast<uint32_t>(frame->width);
result.height = static_cast<uint32_t>(frame->height);
result.crop_left = static_cast<uint32_t>(frame->crop_left);
result.crop_right = static_cast<uint32_t>(frame->crop_right);
result.crop_top = static_cast<uint32_t>(frame->crop_top);
result.crop_bottom = static_cast<uint32_t>(frame->crop_bottom);
result.profile = m_codec->profile > 0 ? static_cast<uint32_t>(m_codec->profile) : 0;
result.level = m_codec->level > 0 ? static_cast<uint32_t>(m_codec->level) : 0;
result.sar_width =
frame->sample_aspect_ratio.num > 0
? static_cast<uint16_t>(std::min(frame->sample_aspect_ratio.num, 65535))
: 0;
result.sar_height =
frame->sample_aspect_ratio.den > 0
? static_cast<uint16_t>(std::min(frame->sample_aspect_ratio.den, 65535))
: 0;
result.color_range = static_cast<uint8_t>(frame->color_range);
result.color_primaries = static_cast<uint8_t>(frame->color_primaries);
result.color_trc = static_cast<uint8_t>(frame->color_trc);
result.color_space = static_cast<uint8_t>(frame->colorspace);
result.key_frame = (frame->flags & AV_FRAME_FLAG_KEY) != 0;
return result;
}
[[nodiscard]] Result CopyFrame(const AVFrame* frame, const FrameBuffer& frame_buffer,
Output* output) {
if (frame->width <= 0 || frame->height <= 0) {
return Result::ApiFail;
}
if ((m_config.max_width > 0 && frame->width > m_config.max_width) ||
(m_config.max_height > 0 && frame->height > m_config.max_height)) {
return Result::OversizeDecode;
}
const auto width = static_cast<uint32_t>(frame->width);
const auto height = static_cast<uint32_t>(frame->height);
const auto pitch = AlignUp(width, 256);
const auto chroma_rows = (static_cast<uint64_t>(height) + 1u) / 2u;
const auto required =
static_cast<uint64_t>(pitch) * height + static_cast<uint64_t>(pitch) * chroma_rows;
if (required > frame_buffer.size) {
return Result::FrameBufferSize;
}
auto* dst = static_cast<uint8_t*>(frame_buffer.data);
std::memset(dst, 0, static_cast<size_t>(required));
if (frame->format == AV_PIX_FMT_NV12) {
for (uint32_t y = 0; y < height; y++) {
std::memcpy(dst + static_cast<size_t>(y) * pitch,
frame->data[0] + static_cast<ptrdiff_t>(y) * frame->linesize[0], width);
}
auto* chroma = dst + static_cast<size_t>(pitch) * height;
for (uint32_t y = 0; y < chroma_rows; y++) {
std::memcpy(chroma + static_cast<size_t>(y) * pitch,
frame->data[1] + static_cast<ptrdiff_t>(y) * frame->linesize[1], width);
}
} else {
m_sws = sws_getCachedContext(m_sws, frame->width, frame->height,
static_cast<AVPixelFormat>(frame->format), frame->width,
frame->height, AV_PIX_FMT_NV12, SWS_FAST_BILINEAR, nullptr,
nullptr, nullptr);
if (m_sws == nullptr) {
return Result::ApiFail;
}
uint8_t* output_planes[4] = {dst, dst + static_cast<size_t>(pitch) * height, nullptr,
nullptr};
int output_strides[4] = {static_cast<int>(pitch), static_cast<int>(pitch), 0, 0};
if (sws_scale(m_sws, frame->data, frame->linesize, 0, frame->height, output_planes,
output_strides) != frame->height) {
return Result::ApiFail;
}
}
output->valid = true;
output->error_frame = (frame->flags & AV_FRAME_FLAG_CORRUPT) != 0;
output->buffer_accepted = true;
output->codec_type = m_config.codec_type;
output->width = width;
output->pitch = pitch;
output->height = height;
output->buffer = frame_buffer.data;
output->buffer_size = frame_buffer.size;
{
std::scoped_lock lock(g_picture_mutex);
g_picture_infos[frame_buffer.data] = {this, MakePictureInfo(frame)};
m_picture_buffers.insert(frame_buffer.data);
}
return Result::Ok;
}
void ClearPictureMetadata() {
std::scoped_lock lock(g_picture_mutex);
for (auto* buffer: m_picture_buffers) {
const auto it = g_picture_infos.find(buffer);
if (it != g_picture_infos.end() && it->second.owner == this) {
g_picture_infos.erase(it);
}
}
m_picture_buffers.clear();
}
Config m_config;
AVCodecContext* m_codec = nullptr;
SwsContext* m_sws = nullptr;
bool m_draining = false;
std::mutex m_mutex;
std::unordered_set<void*> m_picture_buffers;
};
bool IsCodecSupported(uint32_t codec_type) {
return GetAvCodecId(codec_type) != AV_CODEC_ID_NONE;
}
Instance* Create(const Config& config) {
if (!IsCodecSupported(config.codec_type)) {
return nullptr;
}
auto* instance = new Instance(config);
if (!instance->Initialize()) {
delete instance;
return nullptr;
}
return instance;
}
void Destroy(Instance* instance) {
delete instance;
}
uint32_t GetCodecType(const Instance* instance) {
return instance->CodecType();
}
Result Decode(Instance* instance, const Input& input, const FrameBuffer& frame_buffer,
Output* output) {
return instance->DecodeInput(input, frame_buffer, output);
}
Result Flush(Instance* instance, const FrameBuffer& frame_buffer, Output* output) {
return instance->FlushOutput(frame_buffer, output);
}
void Reset(Instance* instance) {
instance->ResetDecoder();
}
bool GetPictureInfo(void* frame_buffer, PictureInfo* picture_info) {
std::scoped_lock lock(g_picture_mutex);
const auto it = g_picture_infos.find(frame_buffer);
if (it == g_picture_infos.end()) {
return false;
}
*picture_info = it->second.info;
return true;
}
} // namespace Libs::VideoDec2::Decoder

Some files were not shown because too many files have changed in this diff Show More