载入中...
搜索中...
未找到
eve::tensor 命名空间参考

命名空间

namespace  affine
 
namespace  glsl_detail
 
namespace  kernels
 
namespace  onnx_detail
 
namespace  q
 
namespace  wgsl_detail
 

类

class  CompiledFunction
 Optimized / scheduled graph ready to run with feeds. 更多...
 
class  Func
 Trace builder — TF2 tf.function analogue (tf.func in scripts). While active, TF ops record into this graph. 更多...
 
struct  FusedGroup
 FusedGroup public API. 更多...
 
class  GpuProgram
 GPU execution of a compiled tensor Graph via generated compute shaders. 更多...
 
class  Graph
 EVENGINE_API_DOMAINS public API. 更多...
 
struct  GraphNode
 GraphNode public API. 更多...
 
struct  KernelSpec
 KernelSpec public API. 更多...
 
struct  OnnxBuffer
 Immutable buffer shared by graph aliases; exactly one storage is authoritative. 更多...
 
class  OnnxCompute
 GPU execution boundary for native ONNX; retains no model and retains compiled resources for the lifetime of the provider. 更多...
 
class  OnnxDeviceStorage
 Run-scoped device storage. Readback completes pending work on the device thread. 更多...
 
struct  OnnxGpuResult
 Owning GPU execution output and completed dispatch count. 更多...
 
struct  OnnxKernel
 Synchronous GPU kernel request; all inputs are borrowed only until dispatch returns. 更多...
 
class  OnnxModel
 Native ONNX import and CPU/GPU execution using tensor kernels, without ONNX Runtime. 更多...
 
struct  OnnxModelInfo
 Owning admission report; unsupported nodes remain inspectable but cannot execute. 更多...
 
struct  OnnxNamedTensor
 Owning named feed/output value. 更多...
 
struct  OnnxRunOptions
 Per-call deterministic RNG and optional strict finite-output diagnostic. 更多...
 
struct  OnnxTensor
 Owning ONNX boundary tensor: row-major little-endian bytes and exact integer shape. 更多...
 
struct  OnnxTransferStats
 Actual transfers and submissions made during one GPU run. 更多...
 
struct  OptimizedGraph
 OptimizedGraph public API. 更多...
 
class  Tensor
 float32 / int32 tensor (rank 1–6), row-major. Eager: owns a buffer. Symbolic: node in a Func graph (no buffer until run). 更多...
 
class  TF
 TF2-like namespace module. Script: tf <- eve.TF(); Default eager; tf.func() traces a graph for compile/run. 更多...
 

枚举

enum class  OpType : uint8_t {
  Placeholder = 0 , Const , Add , Sub ,
  Multiply , Divide , AddScalar , SubScalar ,
  MulScalar , DivScalar , Neg , Abs ,
  Sqrt , Exp , Log , Sin ,
  Cos , Tanh , Relu , Sigmoid ,
  Gelu , Silu , PowScalar , Clamp ,
  MaximumScalar , MinimumScalar , Where , MatMul ,
  Transpose , Permute , Reshape , Flatten ,
  Softmax , LogSoftmax , LayerNorm , RMSNorm ,
  Conv1d , Conv2d , MaxPool2d , AvgPool2d ,
  Embedding , Concat , Slice , ReduceSum ,
  ReduceMean , ReduceMin , ReduceMax , ArgMax ,
  Cast , ScaledDotProductAttention , Resize2d
}
 OpType public API. 更多...
 
enum class  KernelVariant { Default , TiledMatMul }
 Internal choice of generated matrix multiplication implementation. 更多...
 
enum class  OnnxElement : int {
  Float32 = 1 , UInt8 = 2 , Int8 = 3 , Int32 = 6 ,
  Int64 = 7 , Bool = 9
}
 ONNX wire element types; distinct from block-quantized Tensor storage. 更多...
 
enum class  GroupKind : uint8_t {
  Elementwise , MatMul , Conv1d , Conv2d ,
  MaxPool2d , AvgPool2d , Softmax , LayerNorm ,
  RMSNorm , Reduce , ArgMax , Embedding ,
  Concat , Slice , Permute , Sdpa ,
  Resize2d , Alias
}
 GroupKind public API. 更多...
 
enum class  DType : uint8_t {
  Float32 = 0 , Int32 = 1 , Fp16 = 2 , Fp8E4M3 = 3 ,
  Fp4E2M1 = 4 , Int8 = 5 , Int4 = 6
}
 Tensor element types. 更多...
 

函数

bool gpuReduce (const float *data, int size, int op, float &outResult)
 GPU-accelerated reduction for large eager tensors. op: 0 = sum, 1 = min, 2 = max. Returns false (caller should fall back to CPU) when Vulkan/gpgpu isn't available. This compatibility facade preserves the CPU alternative execution contract; new GPU APIs should return a named status enum or Result.
 
bool generateKernel (const Graph &graph, const FusedGroup &group, KernelSpec &out)
 Generate kernel.
 
bool generateMatMulVariant (const Graph &graph, const FusedGroup &group, bool tiled, KernelSpec &out)
 Generate mat mul variant.
 
void generateKernelWgslImpl (const Graph &graph, const FusedGroup &group, KernelSpec &out)
 
Result< KernelSpec > generateWgslKernel (const Graph &graph, const FusedGroup &group, KernelVariant variant=KernelVariant::Default)
 Lower an optimizer-produced group directly to owning WGSL source and dispatch metadata.
 
EVENGINE_API_DOMAINS Result< std::unique_ptr< OnnxCompute > > createOnnxGpuCompute (uint32_t compilerWorkers=4)
 Create a reusable GPU session for the active engine Gpgpu Vulkan device, or an error.
 
OptimizedGraph optimizeGraph (const Graph &graph, int outputNode)
 Optimize graph.
 
int groupKernelCount (const OptimizedGraph &opt)
 Group kernel count.
 
const char * dtypeName (DType dtype)
 Dtype name.
 
bool parseDType (const std::string &name, DType &out)
 Parse d type.
 
 Module_IMPL (TF, new TF())
 

变量

constexpr int kMaxKernelBindings = 8
 

枚举类型说明

◆ DType

enum class eve::tensor::DType : uint8_t
strong

Tensor element types.

Script-visible tensors are float32; int32 tensors are used for index data (argmax outputs, embedding lookups, cast("int32")). Int32 values are stored losslessly as floats for |v| < 2^24, which comfortably covers model vocabularies / sequence lengths / simulation ids used in games.

枚举值
Float32 
Int32 
Fp16 
Fp8E4M3 
Fp4E2M1 
Int8 
Int4 

在文件 Tensor.h 第 24 行定义.

◆ GroupKind

enum class eve::tensor::GroupKind : uint8_t
strong

GroupKind public API.

Kernel group produced by the optimizer (AITemplate-style fusion).

A group is either:

  • an Elementwise chain: multiple graph nodes fused into ONE generated kernel; nodes inside the chain never materialize a buffer;
  • a MatMul / Conv group with an optional bias + elementwise epilogue fused into the same kernel;
  • a single specialized op kernel (softmax, layernorm, attention, ...);
  • an Alias (reshape / flatten / cast): no kernel, output aliases input.
枚举值
Elementwise 
MatMul 
Conv1d 
Conv2d 
MaxPool2d 
AvgPool2d 
Softmax 
LayerNorm 
RMSNorm 
Reduce 
ArgMax 
Embedding 
Concat 
Slice 
Permute 
Sdpa 
Resize2d 
Alias 

在文件 Optimizer.h 第 24 行定义.

◆ KernelVariant

enum class eve::tensor::KernelVariant
strong

Internal choice of generated matrix multiplication implementation.

枚举值
Default 
TiledMatMul 

在文件 KernelGenWgsl.h 第 8 行定义.

◆ OnnxElement

enum class eve::tensor::OnnxElement : int
strong

ONNX wire element types; distinct from block-quantized Tensor storage.

枚举值
Float32 
UInt8 
Int8 
Int32 
Int64 
Bool 

在文件 OnnxModel.h 第 18 行定义.

◆ OpType

enum class eve::tensor::OpType : uint8_t
strong

OpType public API.

枚举值
Placeholder 
Const 
Add 
Sub 
Multiply 
Divide 
AddScalar 
SubScalar 
MulScalar 
DivScalar 
Neg 
Abs 
Sqrt 
Exp 
Log 
Sin 
Cos 
Tanh 
Relu 
Sigmoid 
Gelu 
Silu 
PowScalar 
Clamp 
MaximumScalar 
MinimumScalar 
Where 
MatMul 
Transpose 
Permute 
Reshape 
Flatten 
Softmax 
LogSoftmax 
LayerNorm 
RMSNorm 
Conv1d 
Conv2d 
MaxPool2d 
AvgPool2d 
Embedding 
Concat 
Slice 
ReduceSum 
ReduceMean 
ReduceMin 
ReduceMax 
ArgMax 
Cast 
ScaledDotProductAttention 
Resize2d 

在文件 Graph.h 第 22 行定义.

函数说明

◆ createOnnxGpuCompute()

Result< std::unique_ptr< OnnxCompute > > eve::tensor::createOnnxGpuCompute ( uint32_t  compilerWorkers = 4)

Create a reusable GPU session for the active engine Gpgpu Vulkan device, or an error.

注解
Retain and reuse this owning provider across runGpu calls. Pipelines, immutable initializers and buffer pools persist until provider destruction or Graphics resource retirement. Cache bounds: 1024 shader variants, 128 MiB initializer payload and 512 MiB device buffers. CPU shader compilation uses a bounded queue; pipeline creation and submission remain on the device thread.
参数
compilerWorkersNumber of CPU compiler workers (1..8); default 4. Source is owned by each job. @lifetime Graphics retirement clears device resources before device destruction; either destruction order is supported. A retired session rejects further GPU work; create a new session for a new device. @thread Graphics thread only, no concurrent calls, reentrancy or device destruction during runGpu.

在文件 OnnxGpgpu.cpp 第 347 行定义.

引用了 compilerWorkers, eve::Diagnostic::error(), eve::Failed, gp, eve::InvalidArgument, s, state, token , 以及 eve::Unsupported.

◆ dtypeName()

const char * eve::tensor::dtypeName ( DType  dtype)

Dtype name.

在文件 Tensor.cpp 第 26 行定义.

引用了 Float32, Fp16, Fp4E2M1, Fp8E4M3, Int32, Int4 , 以及 Int8.

被这些函数引用 eve::tensor::Tensor::getDtype().

◆ generateKernel()

EVENGINE_API_DOMAINS bool eve::tensor::generateKernel ( const Graph &  graph,
const FusedGroup &  group,
KernelSpec &  out 
)

Generate kernel.

Generate the specialized GLSL kernel(s) for a fused group. Returns false when the group cannot be lowered (e.g. too many inputs for the fixed 8-binding descriptor layout) — callers use the CPU interpreter. This compatibility facade preserves the optimizer alternate-execution contract; new APIs should return a named status enum or Result.

在文件 KernelGen.cpp 第 921 行定义.

引用了 Alias, ArgMax, AvgPool2d, Concat, Conv1d, Conv2d, Elementwise, Embedding, graph, group, LayerNorm, MatMul, MaxPool2d, Permute, Reduce, Resize2d, RMSNorm, Sdpa, Slice , 以及 Softmax.

◆ generateKernelWgslImpl()

void eve::tensor::generateKernelWgslImpl ( const Graph &  graph,
const FusedGroup &  group,
KernelSpec &  out 
)

◆ generateMatMulVariant()

bool eve::tensor::generateMatMulVariant ( const Graph &  graph,
const FusedGroup &  group,
bool  tiled,
KernelSpec &  out 
)

Generate mat mul variant.

Generate a specific matmul variant for autotuning (tiled=false: thread-per- output naive; tiled=true: shared-memory 16x16 tiles). Only rank-2 matmuls support the tiled variant.

在文件 KernelGen.cpp 第 950 行定义.

引用了 graph, group , 以及 MatMul.

◆ generateWgslKernel()

Result< KernelSpec > eve::tensor::generateWgslKernel ( const Graph &  graph,
const FusedGroup &  group,
KernelVariant  variant = KernelVariant::Default 
)

Lower an optimizer-produced group directly to owning WGSL source and dispatch metadata.

参数
graphBorrowed, immutable graph; must outlive this synchronous call.
groupValid group produced by optimizeGraph for graph.
variantTiledMatMul is supported only for unquantized rank-2 matrix products.
返回
Owning kernel specification, or Unsupported with a code-generation diagnostic.
注解
Reentrant and CPU-only; retains no graph references and creates no GPU resources.

在文件 KernelGenWgsl.cpp 第 530 行定义.

引用了 error, eve::Diagnostic::error(), eve::Result< T >::failure(), generateKernelWgslImpl(), eve::tensor::wgsl_detail::genMatMul(), graph, group, MatMul, eve::tensor::wgsl_detail::specializeInputBindings(), eve::Result< T >::success(), TiledMatMul, eve::Unsupported , 以及 variant.

◆ gpuReduce()

EVENGINE_API_DOMAINS bool eve::tensor::gpuReduce ( const float *  data,
int  size,
int  op,
float &  outResult 
)

GPU-accelerated reduction for large eager tensors. op: 0 = sum, 1 = min, 2 = max. Returns false (caller should fall back to CPU) when Vulkan/gpgpu isn't available. This compatibility facade preserves the CPU alternative execution contract; new GPU APIs should return a named status enum or Result.

在文件 GpuBackend.cpp 第 401 行定义.

引用了 groups, parts, size , 以及 v.

被这些函数引用 eve::tensor::TF::reduceMax(), eve::tensor::TF::reduceMin() , 以及 eve::tensor::TF::reduceSum().

◆ groupKernelCount()

int eve::tensor::groupKernelCount ( const OptimizedGraph &  opt)

Group kernel count.

Number of groups that need GPU kernels (i.e. not Alias).

在文件 Optimizer.cpp 第 498 行定义.

引用了 Alias, count, g , 以及 eve::tensor::OptimizedGraph::groups.

◆ Module_IMPL()

eve::tensor::Module_IMPL ( TF  ,
new   TF() 
)

◆ optimizeGraph()

EVENGINE_API_DOMAINS OptimizedGraph eve::tensor::optimizeGraph ( const Graph &  graph,
int  outputNode 
)

◆ parseDType()

bool eve::tensor::parseDType ( const std::string &  name,
DType &  out 
)

Parse d type.

在文件 Tensor.cpp 第 39 行定义.

引用了 Float32, Fp16, Fp4E2M1, Fp8E4M3, Int32, Int4, Int8 , 以及 name.

被这些函数引用 eve::tensor::TF::cast() , 以及 eve::tensor::TF::quantizeWeight().

变量说明

◆ kMaxKernelBindings

constexpr int eve::tensor::kMaxKernelBindings = 8
constexpr

Max storage bindings available per generated kernel.

在文件 KernelGen.h 第 69 行定义.

被这些函数引用 eve::tensor::glsl_detail::genElementwise() , 以及 eve::tensor::wgsl_detail::genElementwise().