Files
Ahmed Qasid a6d867e4cc fix: avoid topic fallback for non-Latin titles via pragmatic ASCII transliteration (#1526)
# fix: avoid `topic` fallback for non-Latin titles via pragmatic ASCII
transliteration

> **Scope update (in response to review):** this PR is intentionally
broader than its original "Arabic-only" framing. The implementation
changes URL slug generation for **every non-Latin, non-CJK script** that
`slugify` previously stripped — see *Scope* below for the explicit list.
The goal is *not* linguistically correct romanization; it is "avoid
collapsing to `/topic` by producing a usable ASCII slug."

## What this PR is (and isn't)

**Goal:** when a question title contains characters outside Basic Latin
/ Latin Extended / CJK Han, generate a URL slug that is a deterministic
ASCII approximation instead of letting `slugify` strip everything and
falling back to the literal `"topic"`.

**Non-goal:** this is *not* a linguistically correct multi-language
romanizer. The output is a machine-acceptable ASCII slug, not what a
native speaker would choose. For example, `こんにちは` → `konnichiha` (not
the more natural `kon'nichiwa`), `ไทย` → `aithy` (not `thai`). Treat the
slug as an opaque, stable, indexable identifier — the
path-after-`/questions/<id>/` is for SEO and shareability, the canonical
reference is always the ID.

## The bug

Pure non-Latin titles previously got stripped by `slugify.Slugify`, hit
the empty-result fallback in `htmltext.UrlTitle`, and collapsed to the
literal slug `"topic"`. On a live multilingual site, every Arabic / Thai
/ Japanese-hiragana / Korean / Hebrew / Cyrillic question ended up at
`/questions/<id>/topic`.

## The fix

`UrlTitle()` gets a `convertNonLatin` pre-step that mirrors the existing
`convertChinese` pre-step pattern, using
`github.com/mozillazg/go-unidecode` (same author as `go-pinyin` already
in the repo, to minimise new-dep friction).

```
UrlTitle(title)
  → convertChinese(title)        // pre-existing: Han-block → pinyin
  → convertNonLatin(title)       // NEW: detect non-Latin letters → unidecode to ASCII
  → clearEmoji / slugify / url.QueryEscape / cutLongTitle (unchanged)
```

The non-Latin detector skips ASCII, Latin-1 Supplement, Latin
Extended-A/B, and CJK Han. Inputs that hit none of those non-Latin
letter categories short-circuit and return unchanged, so Latin-only and
Chinese-only inputs remain byte-identical (pinned by tests).

## Scope — what scripts are affected

This PR changes behavior for **any** title containing letters in scripts
that `slugify` doesn't handle. Confirmed by tests in
`pkg/htmltext/htmltext_test.go`:

| Script | Example title | Before | After |
| --- | --- | --- | --- |
| Arabic | `كيف حالك` | `topic` | `kyf-hlk` |
| Mixed Latin + Arabic | `مرحبا hello` | `hello` | `mrhb-hello` |
| Thai | `ไทย ไทย` | `topic` | `aithy-aithy` |
| Japanese hiragana | `こんにちは` | `topic` | `konnichiha` |
| Korean | `안녕하세요` | `topic` | `annyeonghaseyo` |
| Hebrew | `שלום עולם` | `topic` | `shlvm-vlm` |
| Cyrillic | `Привет мир` | `topic` | `privet-mir` |

**Unchanged:**

| Case | Behavior |
| --- | --- |
| Pure Latin (`hello world`) | unchanged → `hello-world` |
| Pure Chinese (`这是一个,标题,title`) | unchanged → `zhe-shi-yi-ge-biao-ti`
(pinyin path) |
| Japanese with Han-block kanji (`日本`) | unchanged → `ri-ben` (caught by
pre-existing pinyin path; treated as Chinese reading, not Japanese — a
pre-existing limitation, **not** introduced by this PR) |
| Emoji only (`😂😂😂`) | unchanged → `topic` |
| Empty / whitespace | unchanged → `topic` |

## Transliteration quality — explicit acknowledgement

`go-unidecode` is a generic Unicode → ASCII approximation. It is **not**
a per-language romanization library. Specifically:

- It will pick *one* approximation per codepoint regardless of language
context. `ใ` → `ai` (Thai romanization is `i` or `ai` depending on
standard), `한` → `han`, `語` → `Yu` (Chinese pinyin reading even when
used in Japanese), etc.
- The result is *good enough* to be a stable, URL-safe,
human-recognizable handle, but speakers of the source language will not
consider it "correct."
- It is deterministic, so the same title always produces the same slug —
important since `url_title` is recomputed on every request.

If maintainers prefer to scope this PR more narrowly (e.g. Arabic only,
and reject Thai/Hebrew/Cyrillic/etc.), the detector in
`containsNonLatin` can be tightened to specific Unicode blocks — but
that means the other scripts continue to collapse to `topic`, which is
the bug we're trying to fix. I'd argue the broader fix is preferable to
a piecemeal one, but happy to narrow if you want.

## Live deployment / real-world verification

This patch has been running in production on
**[ask.namasoft.com](https://ask.namasoft.com)** (an Apache Answer
instance we operate) since deployment, built directly from this branch
via `docker compose build`. The site hosts Arabic-language questions, so
the fix exercises the affected code path on every page load.

Sample question URL on the deployed instance:

> `https://ask.namasoft.com/questions/10010000000000115`

The slug in the URL is the transliterated Arabic title rather than
`topic`. No data migration was needed since `url_title` is computed on
every request from `Title` and never persisted (see *Why this is safe to
ship* below).

## Admin-configurable

The transliteration is gated by a package-level `atomic.Bool` (default
**on**, since the current behavior is objectively broken for affected
users):

- `htmltext.SetTransliterateNonLatin(enabled bool)`
- `htmltext.IsTransliterateNonLatinEnabled() bool`

This is deliberately the minimum surface needed to satisfy "the setting
must be readable from `UrlTitle()`". A follow-up PR can add an admin UI
section that calls `SetTransliterateNonLatin` on save and on startup,
without having to re-plumb every `htmltext.UrlTitle` call site through
`context.Context`.

**Default choice — please confirm:** I picked **default-on** because the
existing `topic` behavior is a bug for affected users. If you'd prefer
default-off for strict backward compat on existing installs, flip the
`init()` in `pkg/htmltext/htmltext.go` to `Store(false)` and surface the
toggle as opt-in.

## Why this is safe to ship

- `url_title` is **not** a persisted column. It's not on the `Question`
entity in `internal/entity/question_entity.go`, no migration has ever
added/dropped it, and every call site (`question_service.go`,
`revision_service.go`, `vote_service.go`,
search/report/review/rank/comment services, controllers, repos)
recomputes it from `Title` at response-build time via
`htmltext.UrlTitle(...)`.
- That means the fix is read-only: existing rows light up with correct
slugs on the next request, with no migration and no data rewrite.
- Rollback is just redeploying the prior image; nothing on disk changes.

## Test coverage

`pkg/htmltext/htmltext_test.go`:

- **`TestUrlTitleTable`** — table-driven, one case per affected script
(the full matrix above), plus:
  - `empty` → `topic`
  - `pure latin unchanged` → byte-identical to pre-fix
- `pure chinese unchanged` → byte-identical to pre-fix (pins existing
pinyin behavior)
- `japanese kanji goes through pinyin path unchanged` → documents the
pre-existing Han-block limitation
  - `emoji only falls back to topic` → unchanged
- `long arabic truncates at cutLongTitle boundary` → exercises the
150-byte cap and UTF-8 boundary safety
- **`TestUrlTitleTransliterationToggle`** — with the toggle off,
non-Latin titles collapse to `topic` (pre-fix behavior); with it on,
they transliterate.
- Existing `TestUrlTitle` left untouched.

Test plan for reviewers:

- [ ] `go test ./pkg/htmltext/...` — all pass
- [ ] Visit the live sample URL above and confirm slug is
transliterated, not `topic`
- [ ] Verify Chinese / Latin / emoji-only / empty behavior is
byte-identical to `main` (covered by table tests)

## Out of scope (intentionally)

- No admin UI / site setting plumbing in this PR — see
*Admin-configurable* above. Happy to do the React `Non-Latin Languages
Handling` admin page + `SiteType` + service / controller / migration in
a follow-up if maintainers want it.
- No change to the `"topic"` empty-result fallback.
- No plugin interface for slug generation — mirrored the existing
`convertChinese` pre-step pattern instead.
- No per-language romanization library — this is an explicit non-goal;
see *Transliteration quality* above.

## Issues / discussion

I didn't find an existing upstream issue covering this — happy to be
pointed at one if there is.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: LinkinStars <linkinstar@foxmail.com>
2026-07-10 10:19:39 +08:00

255 lines
6.3 KiB
Go

/*
* Licensed to the Apache Software Foundation (ASF) under one
* or more contributor license agreements. See the NOTICE file
* distributed with this work for additional information
* regarding copyright ownership. The ASF licenses this file
* to you under the Apache License, Version 2.0 (the
* "License"); you may not use this file except in compliance
* with the License. You may obtain a copy of the License at
*
* http://www.apache.org/licenses/LICENSE-2.0
*
* Unless required by applicable law or agreed to in writing,
* software distributed under the License is distributed on an
* "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
* KIND, either express or implied. See the License for the
* specific language governing permissions and limitations
* under the License.
*/
package htmltext
import (
"io"
"net/http"
"net/url"
"regexp"
"strings"
"sync/atomic"
"unicode"
"unicode/utf8"
"github.com/Machiel/slugify"
"github.com/apache/answer/pkg/checker"
"github.com/apache/answer/pkg/converter"
strip "github.com/grokify/html-strip-tags-go"
"github.com/mozillazg/go-pinyin"
"github.com/mozillazg/go-unidecode"
)
var (
reCode = regexp.MustCompile(`(?ism)<(pre)>.*<\/pre>`)
reCodeReplace = "{code...}"
reLink = regexp.MustCompile(`(?ism)<a.*?[^<]>(.*)?<\/a>`)
reLinkReplace = " [$1] "
reSpace = regexp.MustCompile(` +`)
reSpaceReplace = " "
spaceReplacer = strings.NewReplacer(
"\n", " ",
"\r", " ",
"\t", " ",
)
// Without this, pure non-Latin titles (Arabic, Cyrillic, Hebrew, ...) get
// stripped by slugify and collapse to the "topic" fallback. Chinese is
// handled separately by convertChinese.
transliterateNonLatin atomic.Bool
)
func init() {
transliterateNonLatin.Store(true)
}
// SetTransliterateNonLatin toggles non-Latin script transliteration for URL slugs.
func SetTransliterateNonLatin(enabled bool) {
transliterateNonLatin.Store(enabled)
}
// IsTransliterateNonLatinEnabled reports whether non-Latin transliteration is on.
func IsTransliterateNonLatinEnabled() bool {
return transliterateNonLatin.Load()
}
// ClearText clear HTML, get the clear text
func ClearText(html string) string {
if html == "" {
return html
}
html = reCode.ReplaceAllString(html, reCodeReplace)
html = reLink.ReplaceAllString(html, reLinkReplace)
text := spaceReplacer.Replace(strip.StripTags(html))
// replace multiple spaces to one space
return strings.TrimSpace(reSpace.ReplaceAllString(text, reSpaceReplace))
}
func UrlTitle(title string) (text string) {
title = convertChinese(title)
if transliterateNonLatin.Load() {
title = convertNonLatin(title)
}
title = clearEmoji(title)
title = slugify.Slugify(title)
title = url.QueryEscape(title)
title = cutLongTitle(title)
if len(title) == 0 {
title = "topic"
}
return title
}
func clearEmoji(s string) string {
var ret strings.Builder
rs := []rune(s)
for i := range rs {
if len(string(rs[i])) != 4 {
ret.WriteString(string(rs[i]))
}
}
return ret.String()
}
func convertChinese(content string) string {
has := checker.IsChinese(content)
if !has {
return content
}
return strings.Join(pinyin.LazyConvert(content, nil), "-")
}
// Short-circuits on Latin-only / Chinese-only input so existing slugs stay byte-identical.
func convertNonLatin(content string) string {
if !containsNonLatin(content) {
return content
}
return unidecode.Unidecode(content)
}
func containsNonLatin(content string) bool {
for _, r := range content {
switch {
case r < 0x0080: // ASCII
continue
case r >= 0x0080 && r <= 0x024F: // Latin-1 Supplement, Latin Extended-A/B
continue
case unicode.Is(unicode.Han, r): // handled by convertChinese
continue
case unicode.IsLetter(r):
return true
}
}
return false
}
func cutLongTitle(title string) string {
maxBytes := 150
if len(title) <= maxBytes {
return title
}
truncated := title[:maxBytes]
for len(truncated) > 0 && !utf8.ValidString(truncated) {
truncated = truncated[:len(truncated)-1]
}
return truncated
}
// FetchExcerpt return the excerpt from the HTML string
func FetchExcerpt(html, trimMarker string, limit int) (text string) {
return FetchRangedExcerpt(html, trimMarker, 0, limit)
}
// findFirstMatchedWord returns the first matched word and its index
func findFirstMatchedWord(text string, words []string) (string, int) {
if len(text) == 0 || len(words) == 0 {
return "", 0
}
words = converter.UniqueArray(words)
firstWord := ""
firstIndex := len(text)
for _, word := range words {
if idx := strings.Index(text, word); idx != -1 && idx < firstIndex {
firstIndex = idx
firstWord = word
}
}
if firstIndex != len(text) {
return firstWord, firstIndex
}
return "", 0
}
// getRuneRange returns the valid begin and end indexes of the runeText
func getRuneRange(runeText []rune, offset, limit int) (begin, end int) {
runeLen := len(runeText)
limit = min(runeLen, max(0, limit))
begin = min(runeLen, max(0, offset))
end = min(runeLen, begin+limit)
return
}
// FetchRangedExcerpt returns a ranged excerpt from the HTML string.
// Note: offset is a rune index, not a byte index
func FetchRangedExcerpt(html, trimMarker string, offset int, limit int) (text string) {
if len(html) == 0 {
text = html
return
}
runeText := []rune(ClearText(html))
begin, end := getRuneRange(runeText, offset, limit)
text = string(runeText[begin:end])
if begin > 0 {
text = trimMarker + text
}
if end < len(runeText) {
text += trimMarker
}
return
}
// FetchMatchedExcerpt returns the matched excerpt according to the words
func FetchMatchedExcerpt(html string, words []string, trimMarker string, trimLength int) string {
text := ClearText(html)
matchedWord, matchedIndex := findFirstMatchedWord(text, words)
runeIndex := utf8.RuneCountInString(text[0:matchedIndex])
trimLength = max(0, trimLength)
runeOffset := runeIndex - trimLength
runeLimit := trimLength + trimLength + utf8.RuneCountInString(matchedWord)
textRuneCount := utf8.RuneCountInString(text)
if runeOffset+runeLimit > textRuneCount {
// Reserved extra chars before the matched word
runeOffset = textRuneCount - runeLimit
}
return FetchRangedExcerpt(html, trimMarker, runeOffset, runeLimit)
}
func GetPicByUrl(url string) string {
res, err := http.Get(url)
if err != nil {
return ""
}
defer func() {
_ = res.Body.Close()
}()
pix, err := io.ReadAll(res.Body)
if err != nil {
return ""
}
return string(pix)
}