CS AI

Automating product documentation processing

This case study describes a serverless AWS solution for automated processing of technical documents such as datasheets, catalogs, and certifications for a leading Slovak distributor of electrical installation material and industrial automation. The solution accelerated product processing for the B2B catalog, improved data quality, and reduced operating costs.

90 % faster processing

per document, from an original 15–20 minutes
Technologies
  • Amazon S3
  • AWS Lambda
  • SageMaker
AWS Advanced Tier Services Partner
Skener s dotykovým displejom pri automatickom spracovaní dokumentov

The challenge

Distributors of electrical installation and industrial components process tens of thousands of product sheets and technical documents from different manufacturers every year. Many companies still rely on manual data entry, which is slow, inefficient, and prone to errors. Frequent changes to product parameters and pressure to publish new items quickly in B2B portals create a clear need for automation and accurate data.

  • Document volume and variety: hundreds of suppliers and multiple formats (PDF, DOCX, XLS) in different languages.

  • Manual data extraction: 15–20 minutes per document, with a high risk of errors when technical parameters are copied manually.

  • Delayed product publishing: longer time-to-market for new items in the B2B portal.

Project objectives

  • Reduce document processing time by at least 80%.

  • Reduce manual work for the data team by 50%.

  • Achieve extracted data accuracy of ≥ 98%.

  • Implement a scalable pay-per-use architecture integrated with the ERP and B2B system.

The solution

The proposed solution used a serverless AWS architecture with AI/ML and NLP for processing unstructured documents. This approach provided a practical balance of flexibility, reliability, and operational efficiency. The serverless model scales with document volume without requiring infrastructure management, while AWS AI services support accurate processing and straightforward integration with existing ERP and B2B systems.

Key components:

  • Amazon S3: input storage for uploaded documents.

  • Amazon Textract: automated extraction of text, tables, and key-value pairs.

  • Amazon Comprehend + custom NLP (SageMaker): identification and classification of technical parameters such as voltage, power, dimensions, standards, and IP rating.

  • AWS Lambda: workflow orchestration and transformation into structured JSON.

  • Amazon DynamoDB: storage of extracted data and metadata.

  • API Gateway: integration with ERP systems and the B2B portal.

  • Amazon CloudWatch: monitoring of KPIs, latency, errors, and model quality.

Process:

  1. A supplier file is uploaded to S3 in batch or event-driven mode.

  2. Textract extracts text and tables and passes the result to an AWS Lambda function.

  3. The SageMaker NLP model and Amazon Comprehend identify parameters and map them to internal fields.

  4. AWS Lambda completes validation, generates JSON, and stores the record in DynamoDB.

  5. API Gateway enables export to ERP and publishing in the B2B catalog; manual validation is supported by audit logs.

Implementation

Project phases (3–6 months):

  • Analysis and PoC (4 weeks): audit of input documents, definition of the parameter taxonomy, and a PoC using a sample of 500 documents.

  • Development and training (6–10 weeks): deployment of Textract, training of the custom NLP model in SageMaker, development of AWS Lambda functions and integrations.

  • Integration and testing (4–6 weeks): ERP integration, end‑to‑end testing, and security reviews.

  • Deployment and tuning (2–4 weeks): monitoring, user feedback, and model calibration.

Key decisions:

  • Use a serverless AWS solution for scalability and consumption-based pricing.

  • Use a hybrid approach: automated processing with human validation for exceptions.

Results and benefits

The initial evaluation confirmed measurable improvements in operational efficiency, data quality, and product time-to-market. The combination of AI and a serverless approach helped the system process high document volumes while adapting to actual demand and maintaining reliable, scalable performance as workloads increased.

KPI 1: processing speed

  • Baseline: 15–20 min/document.

  • Result: average 1:45 min/document.

  • Impact: approximately 90% faster processing and significantly faster product uploads to the catalog.

KPI 2: data accuracy

  • Baseline: 93–95% manually.

  • Result: 98.6% after validation.

  • Impact: fewer complaints and higher-quality full‑text search.

KPI 3: cost efficiency

  • Result: approximately 40% reduction in data-team costs.

  • Impact: lower OPEX and more resources available for catalog and UX development.

Operational benefits:

  • Scalable parallel processing of thousands of documents.

  • Better data versioning and auditability using DynamoDB and CloudWatch logs.

  • Multilingual support for SK/EN/DE using Amazon Comprehend and custom NLP.

“Automated document processing removed routine work and allowed our team to focus on catalog quality. Faster product uploads directly supported our business objectives.”

Conclusion

The project confirmed that AI-based automation with a serverless architecture can make product documentation processing faster and more accurate in technically demanding sectors. The solution helped the client replace a time-consuming manual process with an agile, measurable system that scales according to demand. The combination of AWS services provides a foundation for extending automation and data management across the organization.

Share

Facing a similar challenge?

Write to us and we will go through what can be done in your environment, with no obligation.