TensorFlow Kernels
In the previous article, we discussed TensorFlow ops. In this article we review kernels.
1. 什么是kernel
如果说op相当于操作的声明,那么kernel就是操作的实现。同一份声明在不同的设备上,最优的实现方式是不一样的,比如对于MatMul矩阵相乘这个操作,在CPU上可以用SSE指令优化加速,在GPU上可以用GPU实现高性能计算。因此就会对应CPU和GPU两种不同的实现。所以,在定义一个kernel时,除了要指明kernel对应的op之外,还需要指明kernel所在的设备类型。
另外,kernel是一个运行期的概念。在定义图时,每个节点的op并不知道具体是由哪个kernel实现这个操作的,因为这时节点还没有被分配到具体的设备上,因此也就没法为其选择合适的kernel。
Kernels
Again, let us start from a real example, the Abs kernel:
1 | REGISTER5(UnaryOp, CPU, "Abs", functor::abs, float, Eigen::half, double, int32, |
where macro REGISTER5 means calling REGISTER for 5 times. And REGISTER calls REGISTER_KERNEL_BUILDER:
1 |
So above call to REGISTER5 expands to
1 | REGISTER_KERNEL_BUILDER(Name("Abs").Device(DEVICE_CPU).TypeConstraint<float>("float"), UnaryOp<CPUDevice, functor::abs<float>>); |
There is another macro invocation for registering GPU versions of Abs:
1 | REGISTER4(UnaryOp, GPU, "Abs", functor::abs, float, Eigen::half, double, int64); |
which expands to:
1 | REGISTER_KERNEL_BUILDER(Name("Abs").Device(DEVICE_GPU).TypeConstraint<float>("float"), UnaryOp<GPUDevice, functor::abs<float>>); |
Each line registers a specialization of class template UnaryOp as a kernel version for a specified device and numeric type.
REGISTER_KERNEL_BUILDER
REGISTER_KERNEL_BUILDER is defined as:
1 |
The introduction of REGISTER_KERNEL_BUILDER_UNIQ_HELPER and REGISTER_KERNEL_BUILDER_UNIQ are to use __COUNTER__, a pre-defined macro provided by GCC and Visual C++:
This macro expands to sequential integral values starting from 0. In conjunction with the ## operator, this provides a convenient means to generate unique identifiers. Care must be taken to ensure that COUNTER is not expanded prior to inclusion of precompiled headers which use it. Otherwise, the precompiled headers will not be used.
From above macro definitions, we see that it’s the constructor of ::tensorflow::kernel_factory::OpKernelRegistrar who registers a kernel. This constructor’s definition is:
1 | typedef OpKernel* (*Factory)(OpKernelConstruction*); |
And one of above variant of Abs expands to
1 | static ::tensorflow::kernel_factory::OpKernelRegistrar registrar__body__4__object( |
KernelDef and KernelDefBuilder
Above Abs example creates the first parameter of type KernelDef by calling the constructor of Name. Name is a sub-class of KernelDefBuilder, whose Build method returns the address of data member KernelDef KernelDefBuilder::kernel_def_. KernelDef is a protobuf message.
KernelDefBuilder‘s constructor fills in KernelDef::op_name.
KernelDefBuilder::Device fills in KernelDef::device_type.
KernelDefBuilder::TypeConstraint is a method template, whose definition is as follows:
1 | template <class T> |
where DataTypeToEnum is a class template, whose each specialization maps a C++ type to an enum ID.
OpKernelRegistrar‘s constructor calls OpKernelRegistrar::InitInternal to register its three parameters to a singleton registry. The third parameter is a C++ lambda, which, if called, allocates and returns a kernel object.
Selective Registration
Note that SHOULD_REGISTER_OP_KERNEL defines a selective registration mechanism like SHOULD_REGISTER_OP does, as we explained in the previous article. SHOULD_REGISTER_OP_KERNEL defaults to true, unless -DSELECTIVE_REGISTRATION is given to GCC, where only classes whose names are listed in variable kNecessaryOpKernelClasses in header file ops_to_register.h would be registered.
Kernel Classes
In above Abs example, the registered kernel class is UnaryOp< CPUDevice, functor::abs<int64>>, where UnaryOp is a sub-class of OpKernel. All kernels are classes derived from OpKernel.
The first template parameter of UnaryOp can take the value of either CPUDevice:
1 | typedef Eigen::ThreadPoolDevice CPUDevice; |
or GPUDevice:
1 | typedef Eigen::GpuDevice GPUDevice; |
The other template parameter of UnaryOp is a functor. In above example, it’s abs:
1 | template <typename T> |
where base defines some types:
1 | template <typename T, typename F, typename R = T> |
These two template parameters are used to implement UnaryOp::Compute:
1 | void Compute(OpKernelContext* ctx) override { |
where UnaryFunctor executes the abs functor:
1 | template <typename Functor> |
Pleae be aware that the Functor::func here refers to base::func, which, in this example, is Eigen::internal::scalar_abs_op<T>. So UnaryOp::Compute actually calls Eigen::internal::scalar_abs_op<T> to compute the abs value.
Template funcion Assign assigns the result to out:
1 | template <typename D, typename Out, typename Rhs> |
Kernel Execution
Note that Compute takes a parameter OpKernelContext* ctx, which provide references to resource managers that allocate resources at graph execution time.
We will review the complete process of kernel execution in a subsequent article.
Kernel Creation
Before executing a kernel, TensorFlow needs to create it by calling the registered factory lambda:
1 | [](::tensorflow::OpKernelConstruction* context) -> ::tensorflow::OpKernel* { |
which in turn calls the constructor of UnaryOp:
1 | template <class T> |
where OP_REQUIRES_OK checks that the data type specified in class template parameter matches the one in OpKernelConstruction parameter passed in by TensorFlow framework at graph creation time.
Conclusion
I will write another article to detail the definition, creation, and execution of a graph.