命名空间 | |
| namespace | detail |
类 | |
| struct | ByteView |
| Borrowed packed 8-bit values, valid for the duration of a synchronous call. 更多... | |
| struct | ConvShape |
| Explicit 1D/2D NCHW convolution geometry; missing 1D height is one. 更多... | |
| struct | QuantizedActivation |
| ONNX unsigned activation quantization; owns bytes and scalar parameters. 更多... | |
函数 | |
| Result< std::vector< uint8_t > > | quantize (std::span< const float > input, float scale, int zeroPoint, bool signedValues) |
| Affine quantization to int8/uint8 bytes, saturating and rounding ties to even. | |
| Result< QuantizedActivation > | dynamicQuantize (std::span< const float > input) |
| Quantize finite FP32 activations with ONNX DynamicQuantizeLinear semantics. | |
| Result< std::vector< float > > | dequantize (ByteView input, std::span< const float > scales, std::span< const int32_t > zeros, size_t inner=1) |
| Affine dequantization using scalar or per-axis scale/zero point. | |
| Result< std::vector< int32_t > > | matmul (ByteView a, ByteView b, size_t m, size_t k, size_t n, int aZero=0, int bZero=0, OnnxCompute *compute=nullptr) |
| Integer row-major [M,K] x [K,N], subtracting scalar zero points. | |
| Result< std::vector< int32_t > > | conv (ByteView x, ByteView w, const ConvShape &shape, int xZero, std::span< const int32_t > wZeros, OnnxCompute *compute=nullptr) |
| Integer Conv, supporting groups, asymmetric padding and dilation. | |
函数说明
◆ conv()
| EVENGINE_API_DOMAINS Result< std::vector< int32_t > > eve::tensor::affine::conv | ( | ByteView | x, |
| ByteView | w, | ||
| const ConvShape & | shape, | ||
| int | xZero, | ||
| std::span< const int32_t > | wZeros, | ||
| OnnxCompute * | compute = nullptr |
||
| ) |
Integer Conv, supporting groups, asymmetric padding and dilation.
- 返回
- Owning NCHW int32 values; rejects bad geometry or accumulator overflow.
- 参数
-
compute Optional borrowed GPU provider; null selects CPU. No retry on GPU failure.
- 注解
- Synchronous. CPU is thread-safe; GPU uses the provider thread. Padding represents real zero.
在文件 AffineQuant.cpp 第 159 行定义.
引用了 b, c, compute, eve::Diagnostic::error(), eve::Failed, eve::tensor::onnx_detail::gpuConv(), eve::InvalidArgument, s, eve::tensor::affine::detail::validateConv(), value, w, x, y , 以及 z.
◆ dequantize()
| Result< std::vector< float > > eve::tensor::affine::dequantize | ( | ByteView | input, |
| std::span< const float > | scales, | ||
| std::span< const int32_t > | zeros, | ||
| size_t | inner = 1 |
||
| ) |
Affine dequantization using scalar or per-axis scale/zero point.
- 参数
-
inner Number of contiguous values following the channel axis.
- 返回
- Owning FP32 values; InvalidArgument on inconsistent channel layout.
- 注解
- CPU, thread-safe; all views are borrowed for the call, no callbacks.
在文件 AffineQuant.cpp 第 62 行定义.
引用了 eve::Diagnostic::error(), input, eve::InvalidArgument, scales, value , 以及 zeros.
◆ dynamicQuantize()
| EVENGINE_API_DOMAINS Result< QuantizedActivation > eve::tensor::affine::dynamicQuantize | ( | std::span< const float > | input | ) |
Quantize finite FP32 activations with ONNX DynamicQuantizeLinear semantics.
- 参数
-
input Borrowed flat input; not retained.
- 返回
- Owning uint8 values, scale, zero point; InvalidArgument for nonfinite input.
- 注解
- CPU, thread-safe, no callbacks; nearest-even rounding independent of fenv.
在文件 AffineQuant.cpp 第 41 行定义.
引用了 bytes, eve::Diagnostic::error(), eve::Result< T >::failure(), input, eve::InvalidArgument, quantize(), eve::tensor::affine::QuantizedActivation::scale, eve::Result< T >::success(), eve::tensor::affine::QuantizedActivation::values, x , 以及 eve::tensor::affine::QuantizedActivation::zeroPoint.
被这些函数引用 eve::tensor::onnx_detail::executeQuant() , 以及 eve::tensor::onnx_detail::executeQuantLstm().
◆ matmul()
| EVENGINE_API_DOMAINS Result< std::vector< int32_t > > eve::tensor::affine::matmul | ( | ByteView | a, |
| ByteView | b, | ||
| size_t | m, | ||
| size_t | k, | ||
| size_t | n, | ||
| int | aZero = 0, |
||
| int | bZero = 0, |
||
| OnnxCompute * | compute = nullptr |
||
| ) |
Integer row-major [M,K] x [K,N], subtracting scalar zero points.
- 返回
- Exact owning int32 accumulators; rejects dimension and int32 overflow.
- 参数
-
compute Optional borrowed GPU provider; null selects CPU. No retry on GPU failure.
- 注解
- Synchronous. CPU is thread-safe; GPU uses the provider thread. Weights remain packed.
在文件 AffineQuant.cpp 第 80 行定义.
引用了 a, b, c, compute, eve::Diagnostic::error(), eve::Failed, eve::tensor::onnx_detail::gpuMatmul(), eve::InvalidArgument, m, n, r, value , 以及 z.
◆ quantize()
| EVENGINE_API_DOMAINS Result< std::vector< uint8_t > > eve::tensor::affine::quantize | ( | std::span< const float > | input, |
| float | scale, | ||
| int | zeroPoint, | ||
| bool | signedValues | ||
| ) |
Affine quantization to int8/uint8 bytes, saturating and rounding ties to even.
- 返回
- Owning packed bytes; rejects invalid scale, zero point or nonfinite input.
- 注解
- CPU, thread-safe; spans are borrowed only during the call.
在文件 AffineQuant.cpp 第 24 行定义.
引用了 eve::Diagnostic::error(), input, eve::InvalidArgument, q, scale , 以及 x.
被这些函数引用 dynamicQuantize() , 以及 eve::tensor::onnx_detail::executeQuant().