命名空间 | |
| namespace | affine |
| namespace | glsl_detail |
| namespace | kernels |
| namespace | onnx_detail |
| namespace | q |
| namespace | wgsl_detail |
类 | |
| class | CompiledFunction |
| Optimized / scheduled graph ready to run with feeds. 更多... | |
| class | Func |
Trace builder — TF2 tf.function analogue (tf.func in scripts). While active, TF ops record into this graph. 更多... | |
| struct | FusedGroup |
| FusedGroup public API. 更多... | |
| class | GpuProgram |
| GPU execution of a compiled tensor Graph via generated compute shaders. 更多... | |
| class | Graph |
| EVENGINE_API_DOMAINS public API. 更多... | |
| struct | GraphNode |
| GraphNode public API. 更多... | |
| struct | KernelSpec |
| KernelSpec public API. 更多... | |
| struct | OnnxBuffer |
| Immutable buffer shared by graph aliases; exactly one storage is authoritative. 更多... | |
| class | OnnxCompute |
| GPU execution boundary for native ONNX; retains no model and retains compiled resources for the lifetime of the provider. 更多... | |
| class | OnnxDeviceStorage |
| Run-scoped device storage. Readback completes pending work on the device thread. 更多... | |
| struct | OnnxGpuResult |
| Owning GPU execution output and completed dispatch count. 更多... | |
| struct | OnnxKernel |
| Synchronous GPU kernel request; all inputs are borrowed only until dispatch returns. 更多... | |
| class | OnnxModel |
| Native ONNX import and CPU/GPU execution using tensor kernels, without ONNX Runtime. 更多... | |
| struct | OnnxModelInfo |
| Owning admission report; unsupported nodes remain inspectable but cannot execute. 更多... | |
| struct | OnnxNamedTensor |
| Owning named feed/output value. 更多... | |
| struct | OnnxRunOptions |
| Per-call deterministic RNG and optional strict finite-output diagnostic. 更多... | |
| struct | OnnxTensor |
| Owning ONNX boundary tensor: row-major little-endian bytes and exact integer shape. 更多... | |
| struct | OnnxTransferStats |
| Actual transfers and submissions made during one GPU run. 更多... | |
| struct | OptimizedGraph |
| OptimizedGraph public API. 更多... | |
| class | Tensor |
| float32 / int32 tensor (rank 1–6), row-major. Eager: owns a buffer. Symbolic: node in a Func graph (no buffer until run). 更多... | |
| class | TF |
TF2-like namespace module. Script: tf <- eve.TF(); Default eager; tf.func() traces a graph for compile/run. 更多... | |
函数 | |
| bool | gpuReduce (const float *data, int size, int op, float &outResult) |
| GPU-accelerated reduction for large eager tensors. op: 0 = sum, 1 = min, 2 = max. Returns false (caller should fall back to CPU) when Vulkan/gpgpu isn't available. This compatibility facade preserves the CPU alternative execution contract; new GPU APIs should return a named status enum or Result. | |
| bool | generateKernel (const Graph &graph, const FusedGroup &group, KernelSpec &out) |
| Generate kernel. | |
| bool | generateMatMulVariant (const Graph &graph, const FusedGroup &group, bool tiled, KernelSpec &out) |
| Generate mat mul variant. | |
| void | generateKernelWgslImpl (const Graph &graph, const FusedGroup &group, KernelSpec &out) |
| Result< KernelSpec > | generateWgslKernel (const Graph &graph, const FusedGroup &group, KernelVariant variant=KernelVariant::Default) |
| Lower an optimizer-produced group directly to owning WGSL source and dispatch metadata. | |
| EVENGINE_API_DOMAINS Result< std::unique_ptr< OnnxCompute > > | createOnnxGpuCompute (uint32_t compilerWorkers=4) |
| Create a reusable GPU session for the active engine Gpgpu Vulkan device, or an error. | |
| OptimizedGraph | optimizeGraph (const Graph &graph, int outputNode) |
| Optimize graph. | |
| int | groupKernelCount (const OptimizedGraph &opt) |
| Group kernel count. | |
| const char * | dtypeName (DType dtype) |
| Dtype name. | |
| bool | parseDType (const std::string &name, DType &out) |
| Parse d type. | |
| Module_IMPL (TF, new TF()) | |
变量 | |
| constexpr int | kMaxKernelBindings = 8 |
枚举类型说明
◆ DType
|
strong |
Tensor element types.
Script-visible tensors are float32; int32 tensors are used for index data (argmax outputs, embedding lookups, cast("int32")). Int32 values are stored losslessly as floats for |v| < 2^24, which comfortably covers model vocabularies / sequence lengths / simulation ids used in games.
| 枚举值 | |
|---|---|
| Float32 | |
| Int32 | |
| Fp16 | |
| Fp8E4M3 | |
| Fp4E2M1 | |
| Int8 | |
| Int4 | |
◆ GroupKind
|
strong |
GroupKind public API.
Kernel group produced by the optimizer (AITemplate-style fusion).
A group is either:
- an Elementwise chain: multiple graph nodes fused into ONE generated kernel; nodes inside the chain never materialize a buffer;
- a MatMul / Conv group with an optional bias + elementwise epilogue fused into the same kernel;
- a single specialized op kernel (softmax, layernorm, attention, ...);
- an Alias (reshape / flatten / cast): no kernel, output aliases input.
| 枚举值 | |
|---|---|
| Elementwise | |
| MatMul | |
| Conv1d | |
| Conv2d | |
| MaxPool2d | |
| AvgPool2d | |
| Softmax | |
| LayerNorm | |
| RMSNorm | |
| Reduce | |
| ArgMax | |
| Embedding | |
| Concat | |
| Slice | |
| Permute | |
| Sdpa | |
| Resize2d | |
| Alias | |
在文件 Optimizer.h 第 24 行定义.
◆ KernelVariant
|
strong |
Internal choice of generated matrix multiplication implementation.
| 枚举值 | |
|---|---|
| Default | |
| TiledMatMul | |
在文件 KernelGenWgsl.h 第 8 行定义.
◆ OnnxElement
|
strong |
ONNX wire element types; distinct from block-quantized Tensor storage.
| 枚举值 | |
|---|---|
| Float32 | |
| UInt8 | |
| Int8 | |
| Int32 | |
| Int64 | |
| Bool | |
在文件 OnnxModel.h 第 18 行定义.
◆ OpType
|
strong |
OpType public API.
函数说明
◆ createOnnxGpuCompute()
| Result< std::unique_ptr< OnnxCompute > > eve::tensor::createOnnxGpuCompute | ( | uint32_t | compilerWorkers = 4 | ) |
Create a reusable GPU session for the active engine Gpgpu Vulkan device, or an error.
- 注解
- Retain and reuse this owning provider across runGpu calls. Pipelines, immutable initializers and buffer pools persist until provider destruction or Graphics resource retirement. Cache bounds: 1024 shader variants, 128 MiB initializer payload and 512 MiB device buffers. CPU shader compilation uses a bounded queue; pipeline creation and submission remain on the device thread.
- 参数
-
compilerWorkers Number of CPU compiler workers (1..8); default 4. Source is owned by each job. @lifetime Graphics retirement clears device resources before device destruction; either destruction order is supported. A retired session rejects further GPU work; create a new session for a new device. @thread Graphics thread only, no concurrent calls, reentrancy or device destruction during runGpu.
在文件 OnnxGpgpu.cpp 第 347 行定义.
引用了 compilerWorkers, eve::Diagnostic::error(), eve::Failed, gp, eve::InvalidArgument, s, state, token , 以及 eve::Unsupported.
◆ dtypeName()
| const char * eve::tensor::dtypeName | ( | DType | dtype | ) |
Dtype name.
在文件 Tensor.cpp 第 26 行定义.
引用了 Float32, Fp16, Fp4E2M1, Fp8E4M3, Int32, Int4 , 以及 Int8.
被这些函数引用 eve::tensor::Tensor::getDtype().
◆ generateKernel()
| EVENGINE_API_DOMAINS bool eve::tensor::generateKernel | ( | const Graph & | graph, |
| const FusedGroup & | group, | ||
| KernelSpec & | out | ||
| ) |
Generate kernel.
Generate the specialized GLSL kernel(s) for a fused group. Returns false when the group cannot be lowered (e.g. too many inputs for the fixed 8-binding descriptor layout) — callers use the CPU interpreter. This compatibility facade preserves the optimizer alternate-execution contract; new APIs should return a named status enum or Result.
在文件 KernelGen.cpp 第 921 行定义.
引用了 Alias, ArgMax, AvgPool2d, Concat, Conv1d, Conv2d, Elementwise, Embedding, graph, group, LayerNorm, MatMul, MaxPool2d, Permute, Reduce, Resize2d, RMSNorm, Sdpa, Slice , 以及 Softmax.
◆ generateKernelWgslImpl()
| void eve::tensor::generateKernelWgslImpl | ( | const Graph & | graph, |
| const FusedGroup & | group, | ||
| KernelSpec & | out | ||
| ) |
◆ generateMatMulVariant()
| bool eve::tensor::generateMatMulVariant | ( | const Graph & | graph, |
| const FusedGroup & | group, | ||
| bool | tiled, | ||
| KernelSpec & | out | ||
| ) |
Generate mat mul variant.
Generate a specific matmul variant for autotuning (tiled=false: thread-per- output naive; tiled=true: shared-memory 16x16 tiles). Only rank-2 matmuls support the tiled variant.
在文件 KernelGen.cpp 第 950 行定义.
◆ generateWgslKernel()
| Result< KernelSpec > eve::tensor::generateWgslKernel | ( | const Graph & | graph, |
| const FusedGroup & | group, | ||
| KernelVariant | variant = KernelVariant::Default |
||
| ) |
Lower an optimizer-produced group directly to owning WGSL source and dispatch metadata.
- 参数
-
graph Borrowed, immutable graph; must outlive this synchronous call. group Valid group produced by optimizeGraph for graph. variant TiledMatMul is supported only for unquantized rank-2 matrix products.
- 返回
- Owning kernel specification, or Unsupported with a code-generation diagnostic.
- 注解
- Reentrant and CPU-only; retains no graph references and creates no GPU resources.
在文件 KernelGenWgsl.cpp 第 530 行定义.
引用了 error, eve::Diagnostic::error(), eve::Result< T >::failure(), generateKernelWgslImpl(), eve::tensor::wgsl_detail::genMatMul(), graph, group, MatMul, eve::tensor::wgsl_detail::specializeInputBindings(), eve::Result< T >::success(), TiledMatMul, eve::Unsupported , 以及 variant.
◆ gpuReduce()
| EVENGINE_API_DOMAINS bool eve::tensor::gpuReduce | ( | const float * | data, |
| int | size, | ||
| int | op, | ||
| float & | outResult | ||
| ) |
GPU-accelerated reduction for large eager tensors. op: 0 = sum, 1 = min, 2 = max. Returns false (caller should fall back to CPU) when Vulkan/gpgpu isn't available. This compatibility facade preserves the CPU alternative execution contract; new GPU APIs should return a named status enum or Result.
在文件 GpuBackend.cpp 第 401 行定义.
引用了 groups, parts, size , 以及 v.
被这些函数引用 eve::tensor::TF::reduceMax(), eve::tensor::TF::reduceMin() , 以及 eve::tensor::TF::reduceSum().
◆ groupKernelCount()
| int eve::tensor::groupKernelCount | ( | const OptimizedGraph & | opt | ) |
Group kernel count.
Number of groups that need GPU kernels (i.e. not Alias).
在文件 Optimizer.cpp 第 498 行定义.
引用了 Alias, count, g , 以及 eve::tensor::OptimizedGraph::groups.
◆ Module_IMPL()
| eve::tensor::Module_IMPL | ( | TF | , |
| new | TF() | ||
| ) |
◆ optimizeGraph()
| EVENGINE_API_DOMAINS OptimizedGraph eve::tensor::optimizeGraph | ( | const Graph & | graph, |
| int | outputNode | ||
| ) |
Optimize graph.
Optimize a traced graph:
- dead-code elimination + topological sort;
- constant folding of elementwise chains over Const nodes;
- elementwise chain fusion into FusedGroups;
- matmul/conv bias + activation epilogue fusion;
- static memory planning with buffer reuse.
在文件 Optimizer.cpp 第 175 行定义.
引用了 a, Add, Alias, b, best, eve::tensor::FusedGroup::biasNode, c, capacity, Cast, Const, Conv1d, Conv2d, Elementwise, eve::tensor::FusedGroup::epilogue, Flatten, g, graph, eve::tensor::OptimizedGraph::groupOrder, groups, eve::tensor::OptimizedGraph::groups, id, idx, inputs, eve::tensor::FusedGroup::inputs, eve::tensor::kernels::isElementwiseOp(), eve::tensor::FusedGroup::kind, m, MatMul, n, nodeId, eve::tensor::FusedGroup::nodes, eve::tensor::OptimizedGraph::nodeSlot, ok, order, eve::tensor::OptimizedGraph::order, eve::tensor::FusedGroup::outputNode, eve::tensor::OptimizedGraph::outputNode, p, eve::tensor::OptimizedGraph::persistentSlots, Placeholder, Reshape, s, size, slots, eve::tensor::OptimizedGraph::slotSize , 以及 u.
◆ parseDType()
| bool eve::tensor::parseDType | ( | const std::string & | name, |
| DType & | out | ||
| ) |
Parse d type.
在文件 Tensor.cpp 第 39 行定义.
引用了 Float32, Fp16, Fp4E2M1, Fp8E4M3, Int32, Int4, Int8 , 以及 name.
被这些函数引用 eve::tensor::TF::cast() , 以及 eve::tensor::TF::quantizeWeight().
变量说明
◆ kMaxKernelBindings
|
constexpr |
Max storage bindings available per generated kernel.
在文件 KernelGen.h 第 69 行定义.
被这些函数引用 eve::tensor::glsl_detail::genElementwise() , 以及 eve::tensor::wgsl_detail::genElementwise().