Step Edge

Step Edge

by StepFun · Current; officially announced as an edge-deployment base model

Step Edge is StepFun’s text-and-vision foundation model for phones, vehicles, and other edge devices. It supports image and video understanding, GUI grounding, OCR, spatial reasoning, tool use, and local agent workflows. The model is designed for low-latency deployment through StepFun’s NPU engine, but its reported limits and performance figures come from a specific evaluation configuration. No public per-token price or general hosted API specification is provided.

Text Reasoning Coding
Step Edge is the base model in StepFun’s 1+N edge-model architecture. Unlike a conventional cloud-only chat model, it is designed to process text and visual inputs on local hardware such as smartphones, vehicles, and other edge devices. Its focus is practical perception and action: understanding screens, locating interface elements, reasoning about space, analyzing images and video, and supporting multi-turn tool workflows.
Outputs

What Step Edge can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Step Edge
Model type Multimodal
Context window 131K tokens
Maximum output 66K tokens
Release date 2026-07-12
Status Current; officially announced as an edge-deployment base model
Knowledge cutoff notes

StepFun’s official Step Edge announcement does not state a knowledge cutoff for the exact model.

Model notes

Step Edge is the text-and-vision base model in StepFun’s 1+N edge architecture. Step Edge Audio, Step Edge GUI, and Step Edge Gen are separate specialized models. The 131,072 context length and 65,536 maximum-token generation setting are reported in StepFun’s internal vLLM evaluation configuration and may not represent universal public API limits. StepFun reports approximately 4.33 seconds for a 1,024-token text-input workflow and 5.61 seconds for understanding a 768-pixel image under its NPU deployment setup. The official page does not publish a per-token price, knowledge cutoff, fine-tuning specification, or general public model identifier for the exact model.

Model guide

Step Edge: StepFun’s On-Device Vision-Language Model for Local Agents

Step Edge is StepFun’s text-and-vision foundation model for phones, vehicles, and other edge devices. It is designed for local image and screen understanding, GUI grounding, spatial and video reasoning, tool use, and low-latency agent workflows where privacy, responsiveness, or weak connectivity matter.

What is Step Edge?

Step Edge is a multimodal foundation model from StepFun, the China-based AI company Shanghai StepFun Intelligence Co., Ltd. The model combines text and visual understanding for applications that run close to the user or device rather than relying exclusively on a remote cloud service.

Its intended environments include smartphones, automotive systems, local assistants, embedded hardware, and other edge deployments. In this context, “edge” means that inference—the process of running the trained model to produce an answer—can take place on local hardware. This can reduce dependence on network connectivity and may improve responsiveness or privacy for certain workflows.

Step Edge is not primarily presented as a general-purpose hosted chat product with a public per-token API. StepFun positions it as the base model in a broader 1+N architecture for edge intelligence. The model handles core text-and-vision reasoning, while separate specialized models extend the system into audio, GUI action, and image generation.

Where Step Edge fits in StepFun’s lineup

Within StepFun’s 1+N design, Step Edge is the “1”: the central foundation model. The “N” represents specialized models for particular edge tasks. Step Edge Audio is associated with audio understanding and automatic speech recognition, Step Edge GUI focuses on interface understanding and action execution, and Step Edge Gen is intended for local image generation and editing.

These related models are important for understanding the product boundary. Step Edge itself should not be described as an image-generation, speech-generation, or music model merely because the wider architecture includes those capabilities. The supplied specifications identify Step Edge as a text-output model with text, image, and video input support. Image, audio, and other generation functions belong to separate components unless StepFun documents them as part of the base model.

Core capabilities and supported inputs

Step Edge accepts text and visual information. Its documented input profile includes text, images, and video. It does not have a documented audio-input capability in the supplied model specification, although Step Edge Audio exists as a separate specialized model.

  • Image understanding: The model can interpret visual content and extract useful information from images.
  • Screen understanding: It can analyze screenshots and other interface displays.
  • GUI grounding: Grounding connects a description or instruction to a specific visual element, such as a button, field, icon, or menu item.
  • OCR and visual extraction: It can identify text and other information contained in visual inputs.
  • Spatial reasoning: It can reason about positions and relationships between objects or interface elements.
  • Video understanding: It can process visual information across video inputs rather than treating a task as a single still image.
  • Tool use: It supports tool calling and multi-turn workflows in which the model can interpret results and continue an interaction.

These capabilities make Step Edge more suitable for visual agents than a text-only language model. For example, an assistant could inspect a device screen, identify the location of a control, use an available tool, and then interpret the updated screen. The model’s role is understanding and reasoning; the exact ability to execute actions depends on the surrounding GUI, tool, operating system, and deployment integration.

Context and output limits

StepFun reports a maximum model length of 131,072 tokens and a maximum generation setting of 65,536 tokens in its documented evaluation configuration. A token is a unit of text processed by a language model; it may represent a word, part of a word, punctuation, or another small text segment.

These figures are best treated as documented deployment settings rather than guaranteed universal limits for every use of Step Edge. The research identifies them in StepFun’s internal vLLM evaluation configuration, and no general public API specification for the exact model is supplied. Actual limits may depend on the hardware, runtime, quantization, serving configuration, input modality, and integration.

The model produces text output. The supplied specification does not identify direct image, video, audio, music, embedding, or speech output from Step Edge. Tool or action workflows may result in external actions through connected systems, but that is different from the model directly generating a non-text media file.

Local deployment and performance

Step Edge is optimized for local inference using StepFun’s Step Inference NPU engine. An NPU is a processor designed to accelerate neural-network workloads. Running the model through an NPU can make edge deployment more practical than relying only on a general-purpose CPU, although results depend heavily on the device and software stack.

StepFun reports approximately 4.33 seconds of end-to-end latency for a 1,024-token text-input workflow and approximately 5.61 seconds for understanding a 768-pixel image in its NPU deployment setup. These are provider-reported measurements, not universal performance guarantees. They were produced under a particular hardware, runtime, input, and serving configuration.

For a real deployment, performance can change with model quantization, image resolution, video length, prompt size, memory bandwidth, NPU implementation, batching, and the number of tools or application steps involved. A compact local model may also be preferable to a larger cloud model when predictable device-side response matters more than maximum reasoning quality.

Reasoning, coding, and tool use

Step Edge’s reasoning profile is visual and operational as much as linguistic. Its documented functions include spatial reasoning, GUI grounding, image and video understanding, and multi-turn tool execution. This makes it relevant to agents that must interpret a changing environment instead of answering a single isolated question.

Examples include locating an interface control from a natural-language instruction, identifying objects or text in a camera view, interpreting a vehicle display, or combining visual observations with tool results. The model can support these workflows, but the surrounding application must provide the tools and define the permissions. The supplied research does not document a standardized public function-calling schema or a guaranteed action protocol.

Step Edge can be used for coding assistance in the broad sense that its text-and-vision capabilities can support software or automation tasks, particularly when screenshots, interfaces, or local tools are involved. However, the official material does not provide a dedicated coding benchmark, coding-specific context limit, or detailed software-development feature set. It should not automatically be treated as a specialized programming model.

Pricing and API availability

No public per-token input or output price is provided for the exact Step Edge model in the supplied research. The official announcement also does not provide a general-purpose hosted API specification or a public model identifier for Step Edge. Consequently, there is no verified recurring price or usage rate to report.

This distinction matters because a model designed for edge deployment may be evaluated primarily through hardware, licensing, integration, or platform arrangements rather than through a conventional public cloud API price. Developers should confirm current access terms, supported runtimes, hardware requirements, and commercial conditions directly with StepFun before planning a production deployment.

Strengths and limitations

Step Edge’s main strength is the combination of visual understanding and edge-oriented deployment. It is designed for situations where a system must interpret screens, images, video, or spatial relationships while keeping latency and network dependence under control. Local processing can also be useful for privacy-sensitive environments, although local execution alone does not establish a complete privacy guarantee.

Its role as a foundation model also gives it a clear division of labor within StepFun’s 1+N architecture. Developers can use Step Edge for core text-and-vision reasoning and add specialized models when they need speech or image generation.

The limitations are equally important. StepFun does not publish a public per-token price, universal API limit, knowledge cutoff, fine-tuning policy, or comprehensive lifecycle schedule for the exact model in the supplied materials. The reported latency figures come from internal testing and may not transfer to another device. In addition, dedicated audio, GUI-action, and image-generation capabilities are associated with separate models rather than being verified as native outputs of Step Edge.

There is also a practical deployment risk: edge AI performance depends on hardware and integration quality. A model may have a large documented context setting, but an embedded device could still be constrained by memory, thermal limits, power consumption, or application-level timeouts.

When to choose Step Edge

Step Edge is a strong candidate when the central requirement is local multimodal perception rather than the lowest cloud price or the broadest public API ecosystem. Consider it for:

  • On-device assistants that need to understand images or screens.
  • Automotive interfaces and other embedded visual-interaction systems.
  • GUI grounding, screen automation, and visual agents.
  • Applications operating with intermittent or weak network connectivity.
  • Low-latency workflows where local inference is more important than access to a large hosted model.
  • Systems that need text-and-vision reasoning as a base for additional specialized edge models.

Another option may be more appropriate when the application requires a documented hosted API, transparent per-token pricing, mature public tooling, guaranteed audio input, direct image generation, speech generation, or a clearly published fine-tuning and lifecycle policy. Step Edge may also be less suitable if the target device lacks an appropriate NPU or enough resources for the intended context and visual workload.

Bottom line

Step Edge is a specialized on-device text-and-vision foundation model rather than a conventional cloud chat endpoint. Its focus is understanding visual environments, grounding instructions in interfaces, reasoning about space and video, and supporting local tool-driven agents. StepFun’s published evaluation figures suggest a design aimed at practical edge latency, but those results are configuration-dependent.

The model is most compelling for developers building visual assistants, automotive systems, screen-aware automation, and privacy- or connectivity-sensitive agents. Anyone evaluating it for production should separately verify hardware compatibility, access terms, integration interfaces, and whether the required audio, GUI action, or media-generation functions belong to Step Edge itself or to one of StepFun’s separate specialized models.


Answers to Frequently Asked Questions

Is Step Edge an image-generation, speech, or GUI-action model?
Step Edge is primarily a text-output model for text-and-vision reasoning. Image generation, audio understanding, and dedicated GUI action capabilities are associated with separate models in StepFun’s 1+N architecture, including Step Edge Gen, Step Edge Audio, and Step Edge GUI.
Does Step Edge have a public API or per-token pricing?
The supplied information does not provide a public per-token price, general-purpose hosted API specification, or public model identifier for the exact Step Edge model. Developers should confirm access terms, supported runtimes, hardware requirements, and commercial conditions directly with StepFun.
Can Step Edge run locally on a device?
Yes. Step Edge is optimized for local inference through StepFun’s Step Inference NPU engine. Local deployment can reduce reliance on network connectivity and may improve responsiveness or privacy, but performance depends on the device, runtime, quantization, memory, and integration.
What is Step Edge?
Step Edge is a multimodal foundation model from StepFun designed for local or edge deployment. It combines text and visual understanding for smartphones, automotive systems, embedded hardware, local assistants, and visual agents.
What inputs and capabilities does Step Edge support?
Step Edge supports text, images, and video as inputs. Its capabilities include image and screen understanding, OCR, GUI grounding, spatial reasoning, video understanding, tool calling, and multi-turn workflows. Audio input is associated with the separate Step Edge Audio model.


Sources 1
Provider

About StepFun