View all IP Cores RoCEv2 / RDMA transmit engine

stream2roce
Stream sensor data straight into host memory.

The stream2roce IP core turns an AXI4-Stream into RoCEv2 RDMA writes entirely in FPGA logic, streaming data straight into remote host memory. It uses the same transport that GigE Vision 3.0 adopts for high-speed image streaming, works with any Ethernet MAC, and scales from 1 GbE up to 800 Gbit/s.

Vendor-independent. Efinix, Lattice, AMD/Xilinx, Altera and Microchip.

Sensor MIPI, ADC, stream FPGA stream2roce RoCEv2 RDMA retransmit, ICRC Host memory AXI-Stream Ethernet RDMA write Host CPU never touches the payload
No CPU neededpure FPGA data path for RDMA writes
1G to 800GGMII/RGMII up to multi-lane transceivers
Multi-streamindependent RDMA queue pairs
Standard RoCEv2interoperates with standard RDMA hosts
The core in one sentence

A complete RoCEv2 sender in FPGA logic

stream2roce takes one or more data streams and writes them directly into the memory of a remote host over standard Ethernet. The full RoCEv2 stack runs in hardware: transport headers, reliable-connection retransmission, ICRC and UDP, IP and Ethernet framing. There is no embedded network stack, no sensor-side driver and no CPU in the payload path.

Kernel-bypass

Kernel-bypass by design

Data lands directly in pre-registered host memory via RDMA Write. The receiving CPU is only notified once a full frame or buffer has arrived, not per packet.

Reliability

Reliable, in hardware

Reliable-Connection and Unreliable-Connection modes, with per-QP sequence numbers, ACK and NAK handling and automatic retransmission with a run-time-tunable timeout.

Integration

AXI-Stream interface

The payload interface is a plain AXI4-Stream. Cameras, ADCs, LiDAR, radar or an on-chip processing chain: anything on a stream becomes an RDMA flow.

Specifications

Technical summary

The essentials are below. The full interface specification, register map and integration guide come with an evaluation engagement.

Transport & reliability

  • RoCEv2, RDMA over Converged Ethernet v2, with RDMA Write
  • Reliable-Connection and Unreliable-Connection modes
  • Hardware retransmission with ACK and NAK, plus ICRC data integrity
  • Credit-based flow control

Performance

  • Line rate from 1 Gbit/s up to 800 Gbit/s
  • Multiple parallel streams, each with its own queue pair
  • Parametrizable datapath width

Integration

  • AXI4-Stream payload input, AXI4-Lite control interface
  • Compatible with any standard Ethernet MAC
  • Independent transmit and receive clock domains

Delivery & support

  • Portable VHDL, no vendor primitives locked in
  • Verified with UVVM and VUnit test benches
  • Available for Efinix, Lattice, AMD/Xilinx, Altera and Microchip
Architecture

One clean pipeline, from stream to wire

Each queue pair owns its packetizer and a hardware retransmission buffer. A shared framer chain then merges all streams, adds UDP, IP and Ethernet, computes the RoCE ICRC and hands finished frames to any standard Ethernet MAC. A lightweight receive path parses incoming ACKs to close the reliability loop.

P2L2 stream2roce IP CORE PER QUEUE PAIR (N) Input FIFO optional IB packetizer RoCE hdr, ICRC Retransmit ACK/NAK buffer MUX UDP/IP IPv4, ports ICRC Ethernet MAC framing Ethernet MAC / PHY FPGA vendor IP AXI-Lite config interface Clock-domain crossing AETH parse ACK/NAK s_axis network axi_lite config ACK to retransmit
Figure 1. stream2roce transmit datapath. Solid lines carry the AXI4-Stream payload; dashed lines carry the ACK receive path and configuration. The Ethernet MAC and PHY are provided by the FPGA vendor.

Transmit chain

  • Optional input FIFO buffers a full packet per stream so the pipeline holds line rate. It can be disabled by generic.
  • IB packetizer builds the InfiniBand transport header for the RDMA Write and reserves the ICRC field.
  • Retransmission block stores in-flight packets and replays them on NAK or ACK-timeout. Depth is a compile-time generic, the timeout a run-time register.
  • Shared framer chain merges all streams, then adds UDP, IP and Ethernet and computes the ICRC to form standard RoCEv2 frames.

Receive & control

  • AETH parser extracts ACK, NAK and PSN from the return path to drive retransmission.
  • Independent clock domains for transmit and receive, with clean clock-domain crossing.
  • AXI-Lite control configures every addressing field, such as MAC, IPv4, UDP port and timeout.
  • Credit flow control follows RoCEv2 end-to-end flow control.
Positioning

Built for RDMA-based sensor streaming

GV

GigE Vision 3.0 runs on RoCEv2

GigE Vision 3.0, released in 2025, adopts RoCEv2 and RDMA as its high-speed streaming transport next to classic GVSP. Image data is written straight into host memory at 10 to 400 Gbit/s, with no CPU copy. Because the standard defines its data channel directly on RoCEv2, this core can act as the transmit engine of a GigE Vision 3.0 device.

ALT

An independent alternative

Some GPU vendors offer FPGA sensor bridges that stream RoCEv2 RDMA into GPU memory through their own network cards. stream2roce is an independent alternative for that transmit role. It runs on the FPGA you choose and delivers to a standard RoCEv2 host, with no tie to a single FPGA, network-card or GPU vendor.

AspectP2L2 stream2roceProprietary GPU-vendor bridge
RoleFPGA-native RoCEv2 transmit engineFPGA bridge tied to one vendor's host
TransportRoCEv2 RDMA Write, RC and UC, HW retransmissionRoCEv2 RDMA write
Host CPU in data pathNoneNone
Receiving hostAny standard RoCEv2 host×One vendor's card and GPU
GPU lock-inNone×Single GPU ecosystem
Own and extend the IPLicensed and adaptable×Closed platform

stream2roce targets any standard RoCEv2 receiver, so you keep full control over your FPGA, network-card and host choices.

Where it fits

One transport, many industries

Any application that moves high-rate sensor or streaming data into a host, while keeping the CPU free for processing, is a fit.

Machine vision

GigE Vision 3.0 cameras and frame-grabbers, with RDMA image delivery at 10 to 100 GbE.

Medical & surgical imaging

Endoscopy, ultrasound and digital pathology, with low-latency ingest into GPU and AI hosts.

LiDAR, radar & sensor fusion

Autonomous and ADAS test rigs streaming multiple sensors over a single link.

Test, measurement & DAQ

High-rate ADC and instrumentation data captured straight into server memory.

Edge AI & inference

Feed GPU inference pipelines from a vendor-neutral FPGA front end.

FPGA-to-server offload

Move processed results from acceleration fabric into host RAM without a CPU copy.

Availability

From proven core to your configuration

Core availableNow

Multi-stream RoCEv2 transmit core with hardware retransmission and ICRC. Runs at 1 GbE over GMII or RGMII without a transceiver, and at 10 GbE. Hardware-proven against a standard Linux RDMA host and verified with UVVM and VUnit.

Higher line rates & your FPGAOn request

25, 50 and 100 Gbit/s, and up to 400 and 800 Gbit/s. Ported to Efinix, Lattice, AMD/Xilinx, Altera or Microchip.

Protocol & host layerOn request

GigE Vision 3.0 and GVSP transmit layer, connection management and metadata handshake, host driver support.

Bring RDMA-speed streaming to your product

Tell us your sensor interface, target FPGA and line rate. We will scope a stream2roce configuration and share the detailed interface specification and evaluation terms.

P2L2
P2L2 GmbH
Softwarepark 35
4232 Hagenberg im Mühlkreis
Austria

+43 681 81541388

info@p2l2.com