이미지에서 텍스트 추출하기 (Java) - Aspose.OCR 및 GroupDocs.Parser 사용
현대 Java 애플리케이션에서는 문서 사진을 검색 가능하고 편집 가능한 텍스트로 변환하는 것이 자동화, 규정 준수 및 분석을 위한 핵심 요구 사항입니다. **이미지에서 텍스트 추출하기 (Java)**는 이 가이드가 답변하는 정확한 질문입니다. Aspose.OCR의 고정밀 광학 문자 인식과 GroupDocs.Parser의 강력한 레이아웃 인식 파싱을 연결하는 방법을 배우게 되며, 스트림을 처리하여 웹 서비스, 배치 작업 및 데스크톱 도구에 모두 적합한 솔루션을 구현합니다.
빠른 답변
- OCR을 처리하는 라이브러리는 무엇입니까? Aspose.OCR delivers industry‑leading accuracy for printed text.
- OCR 출력 결과를 파싱하는 구성 요소는 무엇입니까? GroupDocs.Parser turns raw strings into structured tables, forms, and paragraphs.
- 최소 Java 버전은? JDK 8 또는 그 이상.
- 프로덕션에 라이선스가 필요합니까? 평가용으로는 체험판으로 충분하며, 정식 라이선스를 사용하면 워터마크가 제거되고 모든 기능을 사용할 수 있습니다.
- 이미지 스트림을 직접 처리할 수 있습니까? 예—두 API 모두
InputStream을 받아 HTTP 업로드에 적합합니다.
“이미지에서 텍스트 추출”이란?
이미지에서 텍스트를 추출한다는 것은 스캔된 페이지나 영수증 사진과 같은 시각적 문자들을 코드가 검색·인덱싱·변환할 수 있는 일반 Unicode 문자열로 변환하는 것을 의미합니다. OCR 엔진은 픽셀 패턴을 분석하고 글리프 형태를 인식하여 텍스트 표현을 출력합니다.
왜 Aspose.OCR와 GroupDocs.Parser를 결합하나요?
Aspose.OCR와 GroupDocs.Parser를 결합하면 고품질 문자 인식과 강력한 레이아웃 분석을 동시에 얻을 수 있습니다. Aspose.OCR는 이미지에서 원시 텍스트를 추출하고, GroupDocs.Parser는 해당 텍스트를 해석해 표, 양식, 다중 열 구조를 식별하여 후속 처리에 바로 사용할 수 있는 구조화된 형식으로 반환합니다.
- Accuracy: Aspose.OCR delivers industry‑leading recognition rates.
- Flexibility: GroupDocs.Parser can detect tables, form fields, and multi‑column layouts, returning data in JSON or Java objects.
- Stream‑friendly: Both libraries read directly from
InputStream, eliminating temporary files and simplifying cloud‑native deployments.
사전 요구 사항
- Java Development Kit: JDK 8+ installed.
- Maven: Preferred build tool (or manual JAR handling if you prefer).
- Aspose OCR library: Add the JAR to your project classpath.
- GroupDocs.Parser for Java: Include via Maven (see below) or download the JAR.
- Basic Java knowledge: You should be comfortable with streams, exception handling, and collections.
GroupDocs.Parser for Java 설정
Maven 설정
Add the repository and dependency to your pom.xml:
<repositories>
<repository>
<id>repository.groupdocs.com</id>
<name>GroupDocs Repository</name>
<url>https://releases.groupdocs.com/parser/java/</url>
</repository>
</repositories>
<dependencies>
<dependency>
<groupId>com.groupdocs</groupId>
<artifactId>groupdocs-parser</artifactId>
<version>25.5</version>
</dependency>
</dependencies>
직접 다운로드
If you prefer not to use Maven, grab the latest JAR from GroupDocs Releases.
라이선스 획득
A valid license unlocks the full feature set for both Aspose OCR and GroupDocs.Parser. You can start with a free trial or purchase a permanent license from the vendor websites.
기본 초기화 및 설정
- Set the license for Aspose OCR:
TheLicenseclass loads a license file (license.lic) from the classpath and activates all OCR features.
import com.aspose.ocr.License;
// Initialize and set the Aspose OCR license
License license = new License();
license.setLicense("YOUR_LICENSE_PATH/AsposeOcrLicensePath");
- Initialize GroupDocs.Parser:
No extra code is required for basic parsing; the library auto‑detects the OCR output format when you pass the recognized string.
이미지에서 텍스트를 추출하는 방법 (Java)?
Load an image stream, run Aspose.OCR’s recognizePage method, and feed the resulting text into GroupDocs.Parser—all in under a dozen lines of Java. This direct approach eliminates intermediate files and gives you structured results ready for database insertion or search‑engine indexing.recognizePage processes the supplied image and returns the recognized text as a string.
기능: 이미지 스트림에서 텍스트 인식
개요
The process converts the incoming InputStream to a BufferedImage, optionally limits the OCR to a specific region, and calls Aspose OCR’s recognizePage method. The returned string is then handed to GroupDocs.Parser for layout analysis.
단계별 설명
- Create the AsposeOCR instance:
TheOcrEngineclass is the entry point for all recognition tasks. It encapsulates language models, preprocessing filters, and output settings.
import com.aspose.ocr.AsposeOCR;
AsposeOCR api = new AsposeOCR();
- Read the image stream into a BufferedImage:
BufferedImageis a Java class that stores an image in memory with accessible pixel data.ImageIO.readdecodes the byte stream into a raster image that the OCR engine can analyze. Using aBufferedImagealso lets you crop or rotate the picture before recognition.
import java.awt.image.BufferedImage;
import javax.imageio.ImageIO;
BufferedImage image = ImageIO.read(imageStream);
- Configure recognition settings (optional area selection):
You can limit OCR to a rectangle (Rectangleobject) to speed up processing and reduce false positives when you know the region of interest (e.g., a passport MRZ).
import com.aspose.ocr.RecognitionSettings;
RecognitionSettings settings = new RecognitionSettings();
// Example: limit OCR to a specific rectangle
if (options != null && options.getRectangle() != null) {
ArrayList<Rectangle> areas = new ArrayList<>();
areas.add(new Rectangle(
(int) options.getRectangle().getLeft(),
(int) options.getRectangle().getTop(),
(int) options.getRectangle().getSize().getWidth(),
(int) options.getRectangle().getSize().getHeight()));
settings.setRecognitionAreas(areas);
}
- Run the recognition and handle warnings:
TherecognizePagecall returns aRecognitionResultthat contains the extracted text and any diagnostic warnings (e.g., low confidence segments). Checkresult.getWarnings()to log potential quality issues.
import com.aspose.ocr.RecognitionResult;
RecognitionResult result = api.RecognizePage(image, settings);
if (options != null && options.getHandler() != null) {
options.getHandler().onWarnings(pageIndex, result.warnings);
}
return result.recognitionText;
기능: 이미지 스트림에서 텍스트 영역 인식
개요
When you need each block of text separately—such as individual fields on a form—enable area detection. The OCR engine then returns a list of bounding boxes together with their textual content, which GroupDocs.Parser can map to a structured model.
단계별 설명
- Enable area detection:
SettingrecognitionSettings.setDetectAreas(true)instructs the engine to return rectangle coordinates for every detected text snippet.
RecognitionSettings settings = new RecognitionSettings();
settings.setDetectAreas(true);
(Optional) Define specific regions – reuse the rectangle logic from the previous section if you only care about certain parts of the image.
Execute OCR and collect area information:
The result includes a collection ofTextAreaobjects, each exposinggetRectangle()andgetText(). You can iterate over this collection to populate a DTO or JSON payload.
import java.awt.Rectangle;
import java.util.ArrayList;
ArrayList<PageTextArea> areas = new ArrayList<>();
for (int i = 0; i < result.recognitionAreasRectangles.size(); i++) {
Rectangle rect = result.recognitionAreasRectangles.get(i);
String text = result.recognitionText;
areas.add(new PageTextArea(
text,
new Page(pageIndex, pageSize),
new Rectangle(
new Point(rect.getX(), rect.getY()),
new Size(rect.getWidth(), rect.getHeight()))));
}
return areas;
실용적인 적용 사례
- Document management systems: Index scanned PDFs so users can search the full text without opening the original scan.
- Automated data entry: Pull line‑item details from photographed receipts, invoices, or shipping labels.
- Content digitization: Convert printed manuals into searchable e‑books, preserving tables and headings.
- Compliance monitoring: Scan regulatory forms and automatically flag missing or malformed fields.
성능 고려 사항
- Batch processing: Group up to 20 images per JVM thread to amortize OCR model loading overhead.
- Image quality: Scans at 300 dpi or higher improve recognition accuracy by up to 15 % compared with 150 dpi images.
- Memory management: Call
bufferedImage.flush()after each OCR pass and reuse the sameOcrEngineinstance to keep the native model in memory.
일반적인 문제 및 해결 방법
| Symptom | Likely cause | Fix |
|---|---|---|
| Garbled characters | Low‑resolution image | Use a scan of ≥300 dpi; apply image sharpening before OCR |
| No text returned | Unsupported color space (CMYK) | Convert the image to RGB with BufferedImage.TYPE_INT_RGB |
| Out‑of‑memory errors | Very large images (e.g., >10 MP) | Process the image in tiles or increase the JVM heap (-Xmx4g) |
자주 묻는 질문
Q: How do I install Aspose OCR in my Maven project?
A: Add the Aspose OCR dependency from the Aspose Maven repository to your pom.xml and run mvn clean install. The JAR will be resolved automatically.
Q: Can I extract text from multi‑page PDFs?
A: Yes. Convert each PDF page to an image (for example, with Aspose.PDF), then feed each image stream to the OCR method described above.
Q: Does this approach work with handwritten text?
A: Aspose OCR is optimized for printed characters. For handwriting, consider a dedicated handwriting‑recognition service such as Azure Computer Vision or Google Cloud Vision.
Q: Is a license required for production use?
A: A trial license is sufficient for evaluation, but a full license removes watermarks, lifts usage limits, and provides priority support for commercial deployments.
Q: How can I improve accuracy for a specific language?
A: Set the language on the RecognitionSettings object (e.g., settings.setLanguage(Language.Spanish);). This narrows the character set and dictionary, raising confidence scores.
Last Updated: 2026-08-26
Tested With: Aspose.OCR 23.12, GroupDocs.Parser 25.5
Author: Aspose